CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generalization lag of safety alignment in large language models within the code domain, which enables malicious intents to circumvent defenses via legitimate code structures. This work formally defines this failure mode for the first time and proposes a fully automated black-box jailbreak framework based on structured object-oriented code generation. Furthermore, it integrates latent space projection with activation steering to elucidate the underlying mechanisms of this vulnerability. Extensive experiments demonstrate that the proposed approach achieves an attack success rate of 96.25% across eight mainstream commercial large language models, requiring only 1.51 queries on average. These results significantly outperform existing baselines, highlighting both the efficacy of the method and the critical security gaps in current alignment strategies for code-generating models.
📝 Abstract
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
Problem

Research questions and friction points this paper is trying to address.

safety generalization lag
jailbreak attacks
large language models
code completion
safety alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety Generalization Lag
Black-box Jailbreak
Code Completion
Activation Steering
CodeMimicry
Zhen Liang
Zhen Liang
Loughborough University
Intelligent Textile
H
Hai Huang
School of Computer Science and Technology, Zhejiang Sci-Tech University; Zhejiang Key Laboratory of Digital Fashion and Data Governance, Zhejiang Sci-Tech University
Wentao Chen
Wentao Chen
Shanghai Jiao Tong University
Natural Language ProcessingMachine LearningRepresentation Learning