🤖 AI Summary
This study addresses the significant gap between the functionality and security of code generated by large language models (LLMs), investigating whether iterative model updates can bridge this divide. Leveraging the CWEval benchmark—spanning five programming languages and 31 Common Weakness Enumeration (CWE) categories—we conduct a longitudinal analysis of 32 open-source, closed-source, and compact LLMs to systematically examine vulnerability evolution across versions, scales, and languages. As the first work covering multiple open-source and compact model families, we reveal that openness is not a decisive security factor and identify cross-lingual risk disparities for specific CWEs alongside security regressions in newer models. Our findings demonstrate that while model upgrades improve absolute security, the functionality-security gap persists, with compact models exhibiting notably weaker performance. Consequently, we recommend that developers re-evaluate code security following each model iteration.
📝 Abstract
Large Language Models (LLMs) are widely used to generate code. Although their functional plausibility keeps improving, the generated code often contains security vulnerabilities. The functionality-security gap captures code that passes functional tests but fails security tests. A recent longitudinal study of three model families concluded that LLMs become smarter but not safer, with the only considered open-weight family stagnating. Whether this holds for other (open-weight) families and particularly for compact models remains open. We present a longitudinal study of the gap across 32 LLMs from seven model families (five open-weight), covering three successive releases per family in flagship and compact variants. Using CWEval with 119 tasks in five programming languages and 31 CWEs, we compare trajectories across families, model sizes, and languages. Newer models do become safer in absolute terms, although no family closes the gap. Unlike prior work, we find that openness does not separate the families: every considered open-weight family narrows the gap significantly, while Gemini 3.1 Pro keeps a gap as wide as the one reported for Llama. Compact models usually produce less secure code than their flagship counterparts, with notable exceptions (e.g., Gemini 3.7 Flash). At the CWE level, we confirm persistent weaknesses such as log injection (CWE-117) and HTTP response splitting (CWE-113) and regressions in the newest proprietary models on memory and integer weaknesses, and show that the same CWE carries very different risk across languages. From these results, we derive implications for LLM vendors, researchers, and developers. In particular, developers should assume neither that upgrades improve security nor that proprietary models are more secure; they should rerun security checks after each model change and provide secure APIs in the model's context.