Whetstones: Measuring Coevolution Between Adaptive Malware and Behavioral Defense
The literature on self-adapting malware is open-loop: adaptation is evaluated against static detectors in simulators, with fitness computed by experimenters. This report closes the loop. We built both sides of the adaptive-malware confrontation and the selection loop between them, measuring four generations of engagement under pre-registered protocols. On the red side, a trained adaptive adversary completes reconnaissance-persistence-exfiltration task chains with zero detections under four concurrent commercial stacks (Microsoft Defender RTP, Falco, Wazuh 4.9.2, Suricata 8.0.3, 52,151 ET Open rules). On the blue side, a behavioral-grammar detector -- a 0.88M Transformer prior over an 8-token event grammar fused with four further detection faces -- catches that adversary at 93% with a 3.84% onboarding false-positive rate. The payoff is a law chain: blacklisting creates no selection pressure (R1: 12 cells, 0 alerts); grammar-level hunting pushes selection onto the adversary's body (R2: 9 morphs, 0% survival); body constants bound the rhythm gene's reachable space (G0: the clamp backfires); and the fourth generation produced the loop's first fit morph -- H7-full, 100% survival across three runs, task chains complete, zero detections in six patrol rounds -- whose genome realizes the two escape axes theory predicted (cap cession; non-hidden landing). A power-latency frontier prices each escape axis (N* ~ 240 events); a network-side sibling firewall shows the method transfers (85/85 attacks, 0.10% FPR); and three structural asymmetries explain why equilibrium favors the defender pressing the propagation plane. The report contains the full engineering detail: laboratory, both organs, frontier, four generations of rounds, complete evaluation (baselines, OOV hard-subset, an honest ADFA-LD negative result, E1-E4 studies), coevolution economics, and registered protocols E-A through E-K.