parser benchmarking

Designs and implements empirical benchmark suites and measurement protocols to evaluate parsers, including correctness, runtime, memory usage, robustness, and scalability; builds tooling to run standardized experiments across multiple grammars and analyzes performance tradeoffs, variance, and algorithmic behavior.

parserbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

How Should I Build A Benchmark?

Jan 18, 2025
JC
Jialun Cao
🏛️ The Hong Kong University of Science and Technology | The Chinese University of Hong Kong | Sun Yat-Sen University

Existing large language model (LLM) benchmarks for code lack systematic construction guidelines, resulting in poor data quality, incomplete open-sourcing, sample duplication, and sensitive information leakage—severely undermining evaluation validity and reproducibility. Method: We propose How2Bench—the first comprehensive, fine-grained, actionable guideline for code benchmark development across the full lifecycle, comprising 55 evaluation criteria. It is grounded in a benchmark quality analysis framework, a metadata census of 274 code benchmarks published over the past decade, and an empirical human study involving 49 practitioners. Contribution/Results: Our analysis reveals that ~70% of existing benchmarks lack adequate data quality assurance and >10% are not fully open-sourced. How2Bench significantly enhances defect detection capability and promotes a community-wide paradigm shift toward high-quality, transparent, and reproducible benchmark construction.

Large Model EvaluationProgramming Test QualityTransparency and Reproducibility

This study addresses the longstanding trade-off between expressiveness and performance in parsing by systematically evaluating generalized context-free parsers against deterministic baselines. While deterministic parsers such as LL(1) and LR(1) constrain language design, generalized parsers offer greater expressivity but lack comprehensive empirical assessment. The authors implement six generalized algorithms—CYK, Valiant, Earley, GLL, RNGLR, and BRNGLR—in a unified Rust framework and conduct controlled benchmarks across 22 grammars ranging from arithmetic expressions to full C++ and Java specifications. Their rigorous, reproducible analysis reveals that the performance overhead of generalized parsing is substantially lower than commonly assumed: on deterministic grammars, GLR-family parsers are only about three times slower than LR(1) (median), with low variance and high stability, establishing them as the pragmatic choice for real-world applications requiring full context-free expressiveness.

deterministic parsinggeneral context-free parsinggrammar expressiveness

Existing code generation benchmarks lack representativeness of real-world development scenarios, limiting their utility in guiding model deployment and optimization. This work proposes the first multilingual, multitask evaluation benchmark grounded in real developer telemetry data, spanning six programming languages and six canonical task types, with a strong emphasis on ecological validity and absence of data contamination. It introduces a multidimensional evaluation framework combining functional correctness testing, code similarity analysis, and LLM-as-a-judge to enable fine-grained diagnostics and context-aware performance assessment. Systematic evaluation of nine state-of-the-art models reveals substantial disparities in syntactic accuracy, semantic reasoning capabilities, and practical utility, offering actionable empirical insights for model selection and improvement.

benchmarkcode generationdeveloper telemetry

Existing code generation benchmarks suffer from data contamination, inadequate test coverage, and lack of dynamic update mechanisms. To address these issues, this paper introduces CODE2BENCH—a novel end-to-end dynamic evaluation framework. Methodologically, it features: (1) automated task construction via real-time GitHub repository harvesting and scope-graph-based dependency analysis; (2) function-level task categorization and property-based testing (PBT) generation ensuring 100% branch coverage; and (3) a multi-language, contamination-resistant mechanism for continuous benchmark evolution. Built upon this framework, the CODE2BENCH-2505 benchmark comprises 1,163 real-world programming tasks. It is the first to systematically expose critical deficiencies of state-of-the-art LLMs in complex logical reasoning and cross-language transfer capabilities. By establishing rigorous, reproducible, and evolving evaluation protocols, CODE2BENCH sets a new standard for empirical assessment of code generation models.

Addressing data contamination in existing benchmarks for LLMsAutomating rigorous test suite synthesis for functional verificationEvaluating LLMs on real-world code generation tasks effectively

Automating the Analysis of Parsing Algorithms (and other Dynamic Programs)

Dec 29, 2025
TV
Tim Vieira
🏛️ Johns Hopkins University | ETH Zürich

This paper addresses the challenge of establishing performance guarantees for dynamic programming (DP) parsing algorithms in natural language processing. We present the first automated analysis system that unifies program analysis and complexity inference within a DP framework. Our approach integrates static analysis, type inference, abstract interpretation, and dependency graph modeling to enable formal verification and synthesis of efficient data structures. Key contributions include: (1) a unified formal model capturing DP control flow, data flow, and recurrence structure; (2) automatic inference of precise types, detection of dead code, and identification of redundant computations; and (3) generation of tight, parameterized upper bounds on time and space complexity. We evaluate our system on canonical parsing algorithms—including CKY, Earley, and Neural PCFG—demonstrating substantial improvements in both the automation level and precision of complexity analysis.

Automating analysis of parsing algorithms and dynamic programsInferring types, dead code, and verifying algorithm propertiesProviding guarantees on runtime and space complexity bounds

Latest Papers

What's happening recently
View more

Existing benchmarks primarily assess language models on localized programming tasks, failing to capture their capability to construct complete software systems from scratch. This work introduces ProgramBench, the first end-to-end evaluation framework grounded in behavioral equivalence, which requires agents to autonomously design and implement full codebases based solely on program specifications and documentation, with correctness verified through behavioral test suites. The benchmark encompasses 200 real-world software tasks—including CLI tools, FFmpeg, and SQLite—supports unconstrained, open-ended code generation, and incorporates agent-driven fuzz testing to automatically produce behavioral test cases. Evaluation across nine prominent language models reveals that none can fully solve any task; the best-performing model passes 95% of tests on only 3% of tasks and tends to generate single-file implementations structurally divergent from human-written code.

benchmarkingcode generationlanguage models

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

Current code generation evaluation benchmarks overemphasize correctness metrics while neglecting code quality and practical usability. This work proposes a tripartite evaluation framework that integrates complex project-based benchmarking, static code quality analysis, and structured developer reviews, thereby systematically incorporating real-world developer feedback into the assessment of large language models for code generation for the first time. Leveraging a tree-fold evaluation structure and multi-tiered computer science project benchmarks, experiments on GPT-4.1, DeepSeek-V3-0324, and Claude Opus 4 demonstrate that developer reviews effectively uncover critical production-level issues related to maintainability, readability, and engineering conventions, substantially addressing the limitations of traditional correctness-focused evaluations.

code generationcode qualitydeveloper assessment

Existing evaluation benchmarks for large language models in software engineering often suffer from narrow task coverage, single-dimensional metrics, lack of realistic context, and data contamination, limiting their ability to comprehensively assess model robustness, fairness, and practical utility. To address these limitations, this work proposes BEHELM—the first full-stack benchmarking framework tailored for software engineering. BEHELM establishes a unified, standardized, and reproducible evaluation infrastructure through structured software scenario modeling, multi-granularity input-output specifications, and a multidimensional quality metric system encompassing robustness, explainability, fairness, and efficiency. By significantly lowering the barrier to constructing high-quality benchmarks, BEHELM enables systematic cross-task, cross-language, and cross-granularity evaluations, offering the community a more equitable, realistic, and future-oriented assessment paradigm.

benchmarkingdataset contaminationevaluation metrics

This study addresses the limited representativeness of existing code generation benchmarks—such as HumanEval—in covering programming language knowledge, which leads to significant bias in evaluating large language models (LLMs). The authors introduce, for the first time, a systematic approach based on knowledge units (KUs) to quantitatively analyze the coverage gap between benchmarks and real-world projects. They further propose a prompt-engineering-based task synthesis framework that automatically generates 440 new tasks to enhance benchmark representativeness. Experimental results demonstrate that the augmented benchmark achieves substantially improved KU coverage and over 60% better alignment with the knowledge distribution of real projects. Notably, performance of mainstream LLMs drops by 12.54–44.82% on this enhanced benchmark, revealing that prior evaluations have substantially overestimated model capabilities.

benchmark representativenesscode generationempirical study

Hot Scholars

CQ

Ciyang Qing

Department of Linguistics and English Language, University of Edinburgh
KC

Kyunghyun Cho

New York University, Genentech
Machine LearningDeep Learning
SW

Sean Wang

Southern Methodist University, Cox School of Business, Accounting
BiasDiscriminationCapital MarketsInformation Processing
LC

Lu Cheng

Assistant Professor, UIC CS
Socially Responsible AICausal Machine LearningData MiningAI for Good