long-tail benchmarking

Designs and implements benchmark datasets, scenario suites, and closed-loop evaluation protocols that emphasize rare, low-frequency ("long‑tail") cases and semantic or plan‑level scenarios, and analyzes model performance across scenario modes to quantify long‑tail generalization gaps. Uses these benchmarks to identify and characterize failure modes, compare methods, and measure robustness on underrepresented interactions.

long-tailbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing long-horizon benchmarks merely show that agent performance degrades as task length increases, yet they cannot distinguish whether this decline stems from the intrinsic difficulty of extended tasks or from error accumulation across stages. This work introduces the "horizon residual" metric, which quantifies the additional difficulty beyond what is attributable to compounding errors by comparing the actual success rate on full-length tasks against a baseline predicted from short-segment performance. We formally define this concept for the first time and establish a comparable short-task baseline framework incorporating trajectory-induced degradation analysis, context decay modeling, and log-ratio metrics. Our approach emphasizes the necessity of predefined stage segmentation and resource allocation to control confounding variables, providing an attribution tool for long-horizon evaluation and demonstrating that declining aggregate success rates alone are insufficient evidence of length-specific challenges.

agent failurecontext rothorizon residual

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Existing benchmarks struggle to evaluate agents’ ability to maintain and evolve analytical states over extended data science workflows. This work introduces LongDS, a benchmark comprising 68 multi-turn tasks (2,225 interactions in total) derived from real Kaggle notebooks across six domains, which for the first time systematically defines and implements an evaluation framework tailored for long-horizon data science. LongDS incorporates state evolution patterns—such as counterfactual perturbations, rollbacks, and multi-state compositions—with an average dependency span of 11.3 turns. Experiments reveal that state-of-the-art models achieve only 48.45% average accuracy, suffer a nearly 47-percentage-point performance drop in later stages, and exhibit failure rates of 52%–69% attributable to long-horizon reasoning errors, underscoring state maintenance as a core challenge.

agentic data analysisanalytical statebenchmark

This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.

concept bottleneck modelsconcept labelsmodel interpretability

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Dec 09, 2024
AG
Adhiraj Ghosh
🏛️ University of Tübingen | Open-Ψ (Open-Sci) Collective | University of Cambridge

Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.

Aggregating diverse metrics into reliable model scoresEvaluating open-ended capabilities of foundation modelsReducing evaluation cost while maintaining accuracy

Latest Papers

What's happening recently
View more

Existing reinforcement learning agents often overfit to idiosyncratic patterns in closed environments and lack verifiable behavioral generalization. This work proposes the first cross-domain, long-horizon, multi-tool post-training framework, built upon the open-source MoE model Qwen3.5-122B-A10B and combining two-stage supervised fine-tuning (SFT) with reinforcement learning (RL). Training is conducted on 363 tasks across 27 categories within the MCP benchmark, strictly isolating external evaluation tasks and reward signals. Experimental results demonstrate that the proposed approach substantially enhances out-of-distribution transfer performance, achieving consistent gains across five external benchmarks—including Toolathlon (+9.6 percentage points) and τ²-Bench (+5.3 pp)—and even improves performance on SWE-Bench Pro and Terminal-Bench 2 despite the absence of software engineering tasks in training. The study further uncovers four consistent cross-scenario behavioral divergence patterns.

behavioral evaluationcross-benchmark generalizationlong-horizon agents

This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.

AI evaluationbenchmarkingdeployment conditions

This work addresses a critical yet overlooked issue in financial AI agents: despite producing consistent final outputs, their tool invocations and reasoning trajectories often exhibit substantial inconsistencies that are masked when evaluation focuses solely on end results. To tackle this, the authors introduce DFAH-Bench, a novel benchmark grounded in the Determinism-Faithfulness Assurance Harness (DFAH) framework, which formally defines “faithfulness” as the consistency of observable execution across replays. The benchmark enables fine-grained assessment through two metrics—Decision Agreement Rate (DAR) and Tool-path Agreement Rate (TAR). Evaluated on synthetic compliance and financial DataOps datasets across 570 forward-looking cases, the study reveals high decision consistency (94.2–95.1%) but markedly lower consistency in tool paths (66.9–69.4%) and reasoning trajectories (45.0–51.5%), uncovering significant execution-level variability beneath stable outcomes and establishing a new paradigm for reproducible, replay-based evaluation in compliance and DataOps contexts.

agent instabilitydecision reproducibilityfinancial decision-making

Existing agent evaluation benchmarks are limited in task complexity, realism, and domain diversity, making them inadequate for assessing cross-domain, multi-step reasoning and coordination capabilities. This work proposes a high-fidelity, multi-domain customer interaction benchmark comprising 25 progressively challenging real-world scenarios, introducing for the first time highly complex, interwoven cross-domain tasks that substantially enhance compositional depth, interaction richness, and evaluation rigor. Leveraging both automated metrics and human judgment, the study systematically evaluates 12 leading large language models across dimensions including tool use, multi-step reasoning, and dialogue coherence. The project releases open-source data and code, establishing a reproducible and standardized agent evaluation framework that lays the groundwork for future research on agents operating across diverse real-world settings.

agentic systemsbenchmarkingmulti-domain agents

This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.

capability slicedata-evaluation gapevaluation-to-data inference

Hot Scholars

JY

Jiacheng Yang

Nanjing University
🧠 Large Multimodal Models💪 Reinforcement Learning🥽 Visual Reasoning
TP

Thinh Phan

PhD of Computer Science, University of Arkansas
computer vision
WP

Wonpyo Park

Software Engineer at Google
Model OptimizationDeep LearningComputer VisionAI driven drug discovery
HZ

Hao Zhang

Alumni, University of California, Berkeley
Computer VisionMachine Learning