What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

📅 2026-09-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过分析14,767篇论文,探讨了大型语言模型(LLM)评估基准的设计变化及其对模型性能期望的影响。
📝 Abstract
Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?
Problem

Research questions and friction points this paper is trying to address.

large language models
benchmarks
evaluation
performance expectations
research trends
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Benchmark Design
Evaluation Trends
Interaction
Professional Applications
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chao Wang
Independent Researcher.