deterministic long-form benchmarking

Design and implement benchmarks and evaluation procedures that define a single correct long-form target output for each task instance, producing deterministic ground-truth answers suitable for atomic, unit-level correctness checks. Construct the dataset, procedural rules, and scoring methods to enable repeatable deterministic scoring, calibration and ranking metrics, and to minimize label noise through precise answer specifications and evaluation protocols.

deterministiclong-formbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional language model evaluation often conflates the ability to produce assessable responses with the correctness of those responses, thereby masking execution-level failure modes under aggregate accuracy metrics. This work proposes a two-tiered evaluation framework that disentangles scorer-agnostic execution states—such as termination, answer exposure, parseability, and output length—from scorer-dependent correctness judgments. By enforcing a fixed output budget, tracking multidimensional execution trajectories, formulate a verification mechanism driven by coverage auditing, the study systematically uncovers divergent execution behaviors across models on MATH and ARC-Challenge benchmarks. The analysis reveals that extended output lengths can mitigate certain failure modes and demonstrates that verification strategies substantially influence comparative accuracy outcomes.

accuracy metricsbenchmarkingexecution outcomes

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

Although large language models excel at reasoning tasks, they often fail to faithfully execute multi-step procedures specified in prompts. This work introduces a controlled diagnostic benchmark comprising simple arithmetic programs with variable lengths and backtracking dependencies to systematically evaluate execution fidelity across 14 prominent models on 55 datasets. The study reveals characteristic failure modes in long programs—such as step omission, premature answering, erroneous self-correction, under-execution, and hallucinated steps—for the first time. While models achieve a 61% first-answer accuracy on 5-step programs, performance sharply declines to 20% on 95-step programs, underscoring their significant limitations in following complex, extended instructions.

diagnostic benchmarkinstruction followinglarge language models

This work addresses the limitations of existing uncertainty estimation methods for long-form text generation, which struggle to pinpoint token-level errors and rely on noisy human annotations. The authors propose SALT—the first fine-grained evaluation benchmark grounded in deterministic ground truth—spanning six procedurally generated tasks and enabling assessment of correctness, calibration, and confidence ranking from atomic units to full documents. Through systematic analysis of over 50 large language models, the study reveals that confidence ranking consistently fails at the atomic level and identifies two key error sources: prefix contamination propagation and performance degradation induced by extended context lengths. SALT enables fine-grained error detection and risk control without external evaluators, establishing a new paradigm for reliability assessment in long-text generation.

confidence calibrationdeterministic ground truthfine-grained evaluation

Existing evaluations of language models overemphasize final answer accuracy while neglecting the correctness of intermediate reasoning steps—particularly “level-0” process consistency, i.e., faithful execution of elementary, rule-based chains. Method: The authors introduce L0-Bench, a novel benchmark built from synthetically generated Python functions with human-annotated step-by-step execution traces; it formally defines and quantifies level-0 reasoning capability and establishes program execution traces as the gold standard for process correctness evaluation. The benchmark enables multidimensional, controllable assessment—including context length, voting ensemble size, and reasoning step count. Results: Experiments reveal systematic degradation in process consistency across all models as reasoning depth increases; although larger models and inference-augmented variants exhibit greater robustness, they still face fundamental bottlenecks in maintaining basic procedural fidelity. These findings provide critical diagnostic insights and concrete directions for building reliable, stepwise-reasoning systems.

Evaluating procedural correctness in language models via simple program executionImproving level-0 reasoning for more reliable reasoning systemsTesting models' ability to generate step-by-step, error-free execution traces

Latest Papers

What's happening recently
View more

This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.

dataset defectsevaluation reliabilityformal verification

Hot Scholars

MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
MK

Mohammed Khalilia

Staff Applied Scientist - Qualtrics | Adjunct Professor - Birzeit University
Machine LearningComputational HealthNatural Language ProcessingSpeech
TH

Tsung-Han Wu

PhD Student, UC Berkeley
Vision and LanguageComputer VisionActive Learning
YF

Yujia Fu

Beijing University of Posts and Telecommunications
LLMs NLP
DK

Daehwa Kim

Carnegie Mellon University
Human-Computer InteractionSensingInputRobotics