arXiv · 2604.18510
Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks
Abstract
Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes. We compare harmful supervised fine-tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal-feature abliteration in Qwen2.5-7B and Llama-3.1-8B using harmfulness, capability, self-audit, safety reflection, representation similarity, and refusal-direction repair. All three routes reach near-ceiling harmfulness, but SFT causes the broadest capability loss and representational drift; abliteration yields localized, family-dependent refusal-feature suppression; and RLVR largely preserves base-model capability, explicit safety judgments, and representation geometry while retargeting behavior toward compliance. RLVR models consequently remain unusually responsive to safety-reflection prompts despite complying under direct prompting. We further find that harmful RLVR induces capability-blind compliance: models claim to complete unavailable actions and fabricate information about nonexistent entities. Targeted RLVR calibration substantially reduces this false acceptance without restoring safety or degrading general capability. These results show that harmful compliance, harm recognition, and capability awareness are separable behavioral axes, and that self-audit and hallucination patterns are not necessarily robust safety signals under adaptive post-training.
Explore related subjects
Keep this discovery
Md Rysul Kabir, Zoran Tiganj. 2026-04-20. Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks. https://arxiv.org/abs/2604.18510
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.