evaluate model generalization

Design and execute evaluation protocols and benchmarks that measure how models generalize across conditions, domains, modalities, data types, and training regimes, including in‑distribution, out‑of‑distribution, cross‑condition transfer, zero‑shot and few‑shot testing for both single‑modal and multimodal systems. Analyze and compare model rankings, quantify performance gaps and sensitivity to label schemas and prompts, and identify failure modes (including safety/alignment refusal artefacts), producing reproducible metrics and testing procedures for fair comparison.

evaluatemodelgeneralization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey on Large Language Model Benchmarks

Aug 21, 2025
SN
Shiwen Ni
🏛️ Shenzhen Key Laboratory for High Performance Data Mining | Shenzhen Institutes of Advanced Technology | Chinese Academy of Sciences | Southern University of Science and Technology | University of Chinese Academy of Sciences | Institute of Software | University of Science and Technology of China | Shanghai AI Lab | Shanghai University of Electric Power | South China University of Technology | Shenzhen University | Shenzhen MSU-BIT University | Harbin Institute of Technology | Shenzhen University of Advanced

Current LLM benchmarks suffer from pervasive data contamination, cultural-linguistic bias, lack of procedural transparency, and insufficient dynamism, leading to unreliable evaluations. To address this, we conduct the first systematic survey of 283 mainstream LLM benchmarks, proposing a three-dimensional taxonomy—spanning general capabilities, domain-specific competencies, and goal-specific functionalities—that encompasses language understanding, knowledge reasoning, natural sciences, social sciences and humanities, and risk controllability. Through empirical analysis, we uncover structural biases in evaluation objectives, data provenance, and assessment methodologies. Our key contributions are: (1) the first comprehensive benchmark taxonomy map; (2) identification of critical assessment deficiencies; and (3) a novel benchmark design paradigm grounded in trustworthiness, fairness, and adaptability. This work establishes both a theoretical framework and practical guidelines for developing high-fidelity, next-generation LLM evaluation systems.

Addressing unfair cultural and linguistic biases in evaluationAssessing lack of process credibility and dynamic environment testsEvaluating inflated scores from data contamination in benchmarks

Must-Read Papers

Most classic and influential ideas
View more

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

This study addresses critical limitations in current safety evaluations of large language models, which often rely on single-modality API access, one-off executions, and accuracy alone, thereby overlooking key real-world factors such as modality differences, search capabilities, and response consistency. For the first time, it systematically compares the behavior of ChatGPT’s chat interface versus its API—with and without web search enabled—using the BBQ and SafetyBench benchmarks across 401 prompts replicated three times. The evaluation integrates accuracy, consistency, textual similarity, citation grounding, and abstention behavior. Results reveal that the chat interface consistently underperforms the API in accuracy, that enabling search can reduce accuracy by up to 8 percentage points, that 21% of prompts yield inconsistent responses, and that the two modalities differ significantly in citation practices and abstention strategies. These findings challenge accuracy-centric evaluation paradigms and advocate for a multidimensional safety assessment framework better aligned with deployment realities.

citationsconsistencymodality

Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction following, or pairwise preference, and these scores are often used to support claims about deployed alignment. This paper argues that deployment-relevant alignment cannot be inferred from model-level evaluation alone. Alignment claims should instead be indexed to the level at which evidence is collected: model-level, response-level, interaction-level, or deployment-level. Two studies support this position. First, a structured audit of eleven alignment benchmarks, extended to a sixteen-benchmark corpus, dual-coded against an eight-dimension rubric with Cohen's kappa = 0.87, finds that user-facing verification support is absent across every benchmark examined, while process steerability is nearly absent. The few interactional benchmarks identified, including tau-bench, CURATe, Rifts, and Common Ground, remain fragmented in coverage, and benchmark construction rather than data source determines what is measured. Second, a blinded cross-model stress test using 180 transcripts across three frontier models and four scaffolds finds that the same verification scaffold raises one model's verification support to ceiling while leaving another categorically unchanged. This shows that scaffold efficacy is model-dependent and that the gap identified by the audit cannot be closed at the model level alone. We propose a system-level evaluation agenda: alignment profiles instead of single scores, fixed-scaffolding protocols for comparable interactional evaluation, and reporting templates that make the inferential distance between evaluation evidence and deployment claims explicit.

alignment evaluationbenchmark limitationsdeployment-relevant alignment

Benchmarking Transferability: A Framework for Fair and Robust Evaluation

Apr 28, 2025
AK
Alireza Kazemi
🏛️ The University of Queensland

This work addresses the fundamental lack of fairness and robustness in evaluating model transferability across domains. We propose the first systematic, standardized benchmarking framework for assessing cross-domain transfer capability. Our method introduces a unified multi-source-domain–target-domain evaluation protocol, encompassing diverse transfer tasks and perturbation-robustness analysis, and adopts head-training (i.e., linear-probe fine-tuning) as the consistent evaluation paradigm. Empirical analysis reveals significant performance discrepancies among existing transferability metrics under varying experimental settings, undermining their reliability. Our framework substantially improves assessment fidelity, yielding an average 3.5% gain in transfer performance under standard head-training configurations. To foster reproducibility and rigorous comparison, we fully open-source all code, datasets, and evaluation pipelines—establishing a new, standardized paradigm for transferability measurement.

Addressing inconsistencies in transferability measurement methodsEvaluating reliability of transferability scores across domainsProposing standardized framework for robust transferability assessment

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Dec 09, 2024
AG
Adhiraj Ghosh
🏛️ University of Tübingen | Open-Ψ (Open-Sci) Collective | University of Cambridge

Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.

Aggregating diverse metrics into reliable model scoresEvaluating open-ended capabilities of foundation modelsReducing evaluation cost while maintaining accuracy

Latest Papers

What's happening recently
View more

This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.

AI evaluationbenchmarkingdeployment conditions

This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.

behavioral propertiesinterface propertiesmodel verification

Hot Scholars

MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
AP

Alexander Panchenko

Associate Professor for Natural Language Processing
natural language processingword sense disambiguationtext style transferargument mining
JG

Jie Gui

Southeast University, China
Pattern Recognition and Machine LearningArtificial IntelligenceData MiningDeep Learning
JG

Jingcai Guo

Hong Kong Polytechnic University
Efficient AIZero-Shot LearningEdge AIMachine Learning
RG

Robert Geirhos

Research Scientist, Google DeepMind
Understanding DNNsDeep LearningVideoHuman Vision