Score
Designs and implements benchmark datasets, scenario suites, and closed-loop evaluation protocols that emphasize rare, low-frequency ("long‑tail") cases and semantic or plan‑level scenarios, and analyzes model performance across scenario modes to quantify long‑tail generalization gaps. Uses these benchmarks to identify and characterize failure modes, compare methods, and measure robustness on underrepresented interactions.
This study addresses the limitations of large language models in handling low-frequency, domain-specific, and culturally or temporally sensitive long-tail knowledge, a challenge compounded by a lack of systematic understanding of their failure mechanisms. The work proposes the first four-dimensional analytical framework that integrates technical and sociotechnical perspectives. Through a literature review and conceptual modeling, it systematically defines long-tail knowledge, elucidates the mechanisms by which such knowledge is lost or distorted during training and inference, and evaluates how existing mitigation strategies impact fairness, accountability, and user trust. The research further reveals how current evaluation practices obscure long-tail behaviors and identifies critical open challenges in representation—particularly concerning privacy, sustainability, and governance—thereby offering guidance for future research and system design.
Existing long-horizon benchmarks merely show that agent performance degrades as task length increases, yet they cannot distinguish whether this decline stems from the intrinsic difficulty of extended tasks or from error accumulation across stages. This work introduces the "horizon residual" metric, which quantifies the additional difficulty beyond what is attributable to compounding errors by comparing the actual success rate on full-length tasks against a baseline predicted from short-segment performance. We formally define this concept for the first time and establish a comparable short-task baseline framework incorporating trajectory-induced degradation analysis, context decay modeling, and log-ratio metrics. Our approach emphasizes the necessity of predefined stage segmentation and resource allocation to control confounding variables, providing an attribution tool for long-horizon evaluation and demonstrating that declining aggregate success rates alone are insufficient evidence of length-specific challenges.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
Existing benchmarks struggle to evaluate agents’ ability to maintain and evolve analytical states over extended data science workflows. This work introduces LongDS, a benchmark comprising 68 multi-turn tasks (2,225 interactions in total) derived from real Kaggle notebooks across six domains, which for the first time systematically defines and implements an evaluation framework tailored for long-horizon data science. LongDS incorporates state evolution patterns—such as counterfactual perturbations, rollbacks, and multi-state compositions—with an average dependency span of 11.3 turns. Experiments reveal that state-of-the-art models achieve only 48.45% average accuracy, suffer a nearly 47-percentage-point performance drop in later stages, and exhibit failure rates of 52%–69% attributable to long-horizon reasoning errors, underscoring state maintenance as a core challenge.
This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.
Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.
Existing reinforcement learning agents often overfit to idiosyncratic patterns in closed environments and lack verifiable behavioral generalization. This work proposes the first cross-domain, long-horizon, multi-tool post-training framework, built upon the open-source MoE model Qwen3.5-122B-A10B and combining two-stage supervised fine-tuning (SFT) with reinforcement learning (RL). Training is conducted on 363 tasks across 27 categories within the MCP benchmark, strictly isolating external evaluation tasks and reward signals. Experimental results demonstrate that the proposed approach substantially enhances out-of-distribution transfer performance, achieving consistent gains across five external benchmarks—including Toolathlon (+9.6 percentage points) and τ²-Bench (+5.3 pp)—and even improves performance on SWE-Bench Pro and Terminal-Bench 2 despite the absence of software engineering tasks in training. The study further uncovers four consistent cross-scenario behavioral divergence patterns.
This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.
This work addresses a critical yet overlooked issue in financial AI agents: despite producing consistent final outputs, their tool invocations and reasoning trajectories often exhibit substantial inconsistencies that are masked when evaluation focuses solely on end results. To tackle this, the authors introduce DFAH-Bench, a novel benchmark grounded in the Determinism-Faithfulness Assurance Harness (DFAH) framework, which formally defines “faithfulness” as the consistency of observable execution across replays. The benchmark enables fine-grained assessment through two metrics—Decision Agreement Rate (DAR) and Tool-path Agreement Rate (TAR). Evaluated on synthetic compliance and financial DataOps datasets across 570 forward-looking cases, the study reveals high decision consistency (94.2–95.1%) but markedly lower consistency in tool paths (66.9–69.4%) and reasoning trajectories (45.0–51.5%), uncovering significant execution-level variability beneath stable outcomes and establishing a new paradigm for reproducible, replay-based evaluation in compliance and DataOps contexts.
Existing agent evaluation benchmarks are limited in task complexity, realism, and domain diversity, making them inadequate for assessing cross-domain, multi-step reasoning and coordination capabilities. This work proposes a high-fidelity, multi-domain customer interaction benchmark comprising 25 progressively challenging real-world scenarios, introducing for the first time highly complex, interwoven cross-domain tasks that substantially enhance compositional depth, interaction richness, and evaluation rigor. Leveraging both automated metrics and human judgment, the study systematically evaluates 12 leading large language models across dimensions including tool use, multi-step reasoning, and dialogue coherence. The project releases open-source data and code, establishing a reproducible and standardized agent evaluation framework that lays the groundwork for future research on agents operating across diverse real-world settings.
This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.