Complexity-Aware Evaluation of LLM Comprehension

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing aggregated benchmarks, which obscure how the reliability of large language models (LLMs) in code understanding varies with complexity. We propose a complexity-aware evaluation framework that integrates structural metrics—including cyclomatic complexity, nesting depth, branching factor, and Halstead volume—with logistic regression analysis to quantify their impact on model performance. Experimental results demonstrate that while DeepSeek-Coder-V2 achieves an overall accuracy of 78.33%, its performance degrades sharply to 52.78% under high-complexity conditions. By effectively revealing these degradation patterns, the proposed framework offers a more diagnostic evaluation paradigm than traditional aggregated metrics, enabling finer-grained assessment of LLM capabilities across varying levels of code complexity.
📝 Abstract
Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llama through two complementary tasks: automatic input-output prediction over 300 Python functions and manually assessed semantic comprehension over a balanced subset of 60 functions. The functions are grouped into Low-, Medium-, and High-complexity bands. DeepSeek-Coder-V2 achieves an overall automatic accuracy of 78.33%, compared with 70.33% for Llama. However, accuracy decreases substantially from Low to High complexity, from 93.52% to 52.78% for DeepSeek-Coder-V2 and from 87.04% to 47.22% for Llama. Incorrect predictions are consistently associated with higher values of all four complexity metrics, and correlation and logistic-regression analyses confirm broadly comparable negative associations between structural complexity and correctness. Manual semantic comprehension shows the same degradation pattern, with accuracy decreasing from 100.00% to 75.00% for DeepSeek-Coder-V2 and from 90.00% to 60.00% for Llama. These findings demonstrate that complexity-aware evaluation provides a more diagnostic assessment of LLM code-comprehension reliability than aggregate accuracy alone.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Code Comprehension
Structural Complexity
Evaluation Framework
Software Engineering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Complexity-Aware Evaluation
Cyclomatic Complexity
Code Comprehension
Large Language Models
Halstead Volume
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Ali Mohammadi Esfahani
Systems and computer Engineering, Carleton University, Ottawa, Canada
Nafiseh Kahani
Nafiseh Kahani
Carleton University
AI-based System TestingFormal VerificationTrustworthy AI
S
Samuel A. Ajila
Systems and computer Engineering, Carleton University, Ottawa, Canada