Score
Designs and implements empirical benchmark suites and measurement protocols to evaluate parsers, including correctness, runtime, memory usage, robustness, and scalability; builds tooling to run standardized experiments across multiple grammars and analyzes performance tradeoffs, variance, and algorithmic behavior.
Existing large language model (LLM) benchmarks for code lack systematic construction guidelines, resulting in poor data quality, incomplete open-sourcing, sample duplication, and sensitive information leakage—severely undermining evaluation validity and reproducibility. Method: We propose How2Bench—the first comprehensive, fine-grained, actionable guideline for code benchmark development across the full lifecycle, comprising 55 evaluation criteria. It is grounded in a benchmark quality analysis framework, a metadata census of 274 code benchmarks published over the past decade, and an empirical human study involving 49 practitioners. Contribution/Results: Our analysis reveals that ~70% of existing benchmarks lack adequate data quality assurance and >10% are not fully open-sourced. How2Bench significantly enhances defect detection capability and promotes a community-wide paradigm shift toward high-quality, transparent, and reproducible benchmark construction.
This study addresses the longstanding trade-off between expressiveness and performance in parsing by systematically evaluating generalized context-free parsers against deterministic baselines. While deterministic parsers such as LL(1) and LR(1) constrain language design, generalized parsers offer greater expressivity but lack comprehensive empirical assessment. The authors implement six generalized algorithms—CYK, Valiant, Earley, GLL, RNGLR, and BRNGLR—in a unified Rust framework and conduct controlled benchmarks across 22 grammars ranging from arithmetic expressions to full C++ and Java specifications. Their rigorous, reproducible analysis reveals that the performance overhead of generalized parsing is substantially lower than commonly assumed: on deterministic grammars, GLR-family parsers are only about three times slower than LR(1) (median), with low variance and high stability, establishing them as the pragmatic choice for real-world applications requiring full context-free expressiveness.
Existing code generation benchmarks lack representativeness of real-world development scenarios, limiting their utility in guiding model deployment and optimization. This work proposes the first multilingual, multitask evaluation benchmark grounded in real developer telemetry data, spanning six programming languages and six canonical task types, with a strong emphasis on ecological validity and absence of data contamination. It introduces a multidimensional evaluation framework combining functional correctness testing, code similarity analysis, and LLM-as-a-judge to enable fine-grained diagnostics and context-aware performance assessment. Systematic evaluation of nine state-of-the-art models reveals substantial disparities in syntactic accuracy, semantic reasoning capabilities, and practical utility, offering actionable empirical insights for model selection and improvement.
Existing code generation benchmarks suffer from data contamination, inadequate test coverage, and lack of dynamic update mechanisms. To address these issues, this paper introduces CODE2BENCH—a novel end-to-end dynamic evaluation framework. Methodologically, it features: (1) automated task construction via real-time GitHub repository harvesting and scope-graph-based dependency analysis; (2) function-level task categorization and property-based testing (PBT) generation ensuring 100% branch coverage; and (3) a multi-language, contamination-resistant mechanism for continuous benchmark evolution. Built upon this framework, the CODE2BENCH-2505 benchmark comprises 1,163 real-world programming tasks. It is the first to systematically expose critical deficiencies of state-of-the-art LLMs in complex logical reasoning and cross-language transfer capabilities. By establishing rigorous, reproducible, and evolving evaluation protocols, CODE2BENCH sets a new standard for empirical assessment of code generation models.
This paper addresses the challenge of establishing performance guarantees for dynamic programming (DP) parsing algorithms in natural language processing. We present the first automated analysis system that unifies program analysis and complexity inference within a DP framework. Our approach integrates static analysis, type inference, abstract interpretation, and dependency graph modeling to enable formal verification and synthesis of efficient data structures. Key contributions include: (1) a unified formal model capturing DP control flow, data flow, and recurrence structure; (2) automatic inference of precise types, detection of dead code, and identification of redundant computations; and (3) generation of tight, parameterized upper bounds on time and space complexity. We evaluate our system on canonical parsing algorithms—including CKY, Earley, and Neural PCFG—demonstrating substantial improvements in both the automation level and precision of complexity analysis.
Existing benchmarks primarily assess language models on localized programming tasks, failing to capture their capability to construct complete software systems from scratch. This work introduces ProgramBench, the first end-to-end evaluation framework grounded in behavioral equivalence, which requires agents to autonomously design and implement full codebases based solely on program specifications and documentation, with correctness verified through behavioral test suites. The benchmark encompasses 200 real-world software tasks—including CLI tools, FFmpeg, and SQLite—supports unconstrained, open-ended code generation, and incorporates agent-driven fuzz testing to automatically produce behavioral test cases. Evaluation across nine prominent language models reveals that none can fully solve any task; the best-performing model passes 95% of tests on only 3% of tasks and tends to generate single-file implementations structurally divergent from human-written code.
This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.
Current code generation evaluation benchmarks overemphasize correctness metrics while neglecting code quality and practical usability. This work proposes a tripartite evaluation framework that integrates complex project-based benchmarking, static code quality analysis, and structured developer reviews, thereby systematically incorporating real-world developer feedback into the assessment of large language models for code generation for the first time. Leveraging a tree-fold evaluation structure and multi-tiered computer science project benchmarks, experiments on GPT-4.1, DeepSeek-V3-0324, and Claude Opus 4 demonstrate that developer reviews effectively uncover critical production-level issues related to maintainability, readability, and engineering conventions, substantially addressing the limitations of traditional correctness-focused evaluations.
Existing evaluation benchmarks for large language models in software engineering often suffer from narrow task coverage, single-dimensional metrics, lack of realistic context, and data contamination, limiting their ability to comprehensively assess model robustness, fairness, and practical utility. To address these limitations, this work proposes BEHELM—the first full-stack benchmarking framework tailored for software engineering. BEHELM establishes a unified, standardized, and reproducible evaluation infrastructure through structured software scenario modeling, multi-granularity input-output specifications, and a multidimensional quality metric system encompassing robustness, explainability, fairness, and efficiency. By significantly lowering the barrier to constructing high-quality benchmarks, BEHELM enables systematic cross-task, cross-language, and cross-granularity evaluations, offering the community a more equitable, realistic, and future-oriented assessment paradigm.
This study addresses the limited representativeness of existing code generation benchmarks—such as HumanEval—in covering programming language knowledge, which leads to significant bias in evaluating large language models (LLMs). The authors introduce, for the first time, a systematic approach based on knowledge units (KUs) to quantitatively analyze the coverage gap between benchmarks and real-world projects. They further propose a prompt-engineering-based task synthesis framework that automatically generates 440 new tasks to enhance benchmark representativeness. Experimental results demonstrate that the augmented benchmark achieves substantially improved KU coverage and over 60% better alignment with the knowledge distribution of real projects. Notably, performance of mainstream LLMs drops by 12.54–44.82% on this enhanced benchmark, revealing that prior evaluations have substantially overestimated model capabilities.