arXiv · 2609.39902
CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
Abstract
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhen Liang, Hai Huang, Wentao Chen. 2026-09-30. CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion. https://arxiv.org/abs/2609.39902
Cite the original work for its findings. Save a collection to share your selection of sources.