Score
Designs, implements, and analyzes benchmarking suites and evaluation protocols for pretraining large language and vision–language models, including datasets, metrics, baselines, and reproducible training pipelines to compare pretraining objectives, data regimes, architectures, and compute/resource tradeoffs. Measures and reports model properties such as intrinsic natural-language evaluation scores, transfer to downstream tasks, robustness and failure modes, and efficiency so practitioners can objectively compare and improve pretraining methods.
Current LLM benchmarks suffer from pervasive data contamination, cultural-linguistic bias, lack of procedural transparency, and insufficient dynamism, leading to unreliable evaluations. To address this, we conduct the first systematic survey of 283 mainstream LLM benchmarks, proposing a three-dimensional taxonomy—spanning general capabilities, domain-specific competencies, and goal-specific functionalities—that encompasses language understanding, knowledge reasoning, natural sciences, social sciences and humanities, and risk controllability. Through empirical analysis, we uncover structural biases in evaluation objectives, data provenance, and assessment methodologies. Our key contributions are: (1) the first comprehensive benchmark taxonomy map; (2) identification of critical assessment deficiencies; and (3) a novel benchmark design paradigm grounded in trustworthiness, fairness, and adaptability. This work establishes both a theoretical framework and practical guidelines for developing high-fidelity, next-generation LLM evaluation systems.
Existing evaluation methodologies for large language models (LLMs) suffer from insufficient generalization assessment, as static benchmarks fail to capture the continuously expanding capability boundaries of evolving LLMs. Method: We formally define “evaluation generalizability” and propose a four-dimensional analytical framework encompassing evaluation methodologies, datasets, evaluators, and metrics. Our approach innovatively integrates LLM-as-a-judge, dynamically updated datasets, capability-decoupled benchmark design, and a multidimensional meta-evaluation framework. Contribution: We establish a novel, capability-oriented, automated, and sustainably evolvable evaluation paradigm covering critical dimensions—including knowledge, reasoning, instruction following, multimodal understanding, and safety. Concurrently, we release an open-source, extensible GitHub “living review” repository—a community-maintained, versioned resource—to advance evaluation practice from static benchmarking toward dynamic, collaborative co-evolution.
Current LLM pretraining lacks a standardized benchmark for optimizer evaluation, hindering reproducibility and fair comparison. This work introduces the first unified, systematic evaluation framework for optimization algorithms in large-language-model pretraining. We conduct controlled, cross-optimizer comparisons—including AdamW, Lion, and Adafactor—across diverse model scales, batch sizes, and training durations. Crucially, we ensure fairness and reproducibility through rigorous ablation of confounding variables and meticulous hyperparameter tuning. Our analysis uncovers fundamental trade-offs among convergence speed, training stability, and computational efficiency, yielding practical, scenario-aware optimizer selection guidelines. All code, hyperparameter configurations, and experimental results are publicly released to establish a reproducible benchmark and accelerate empirical optimizer research.
Existing language model evaluation benchmarks exhibit substantial ranking inconsistencies—even for similar capabilities—undermining comparability and external validity. To address this, we propose a “train-then-test” paradigm: prior to evaluation, each model undergoes benchmark-specific fine-tuning to standardize assessment conditions. We conduct the first systematic validation across 24 benchmarks and 61 mainstream models. Our results show that this approach significantly improves cross-benchmark ranking consistency—particularly within model families, where rankings become nearly perfectly aligned—reveals latent structural patterns in performance differences, and drives score matrices toward rank-one structure. By integrating benchmark-customized fine-tuning, cross-benchmark correlation analysis, and low-rank modeling, our method substantially enhances evaluation reliability and interpretability. It establishes a more robust, standardized framework for language model assessment.
This work addresses the proliferation of large language model (LLM) evaluation benchmarks, which has outpaced systematic assessment of their intrinsic quality. To this end, we propose Benchmark², a novel framework that establishes the first quantitative methodology for evaluating the reliability and validity of LLM benchmarks through three complementary metrics: cross-benchmark ranking consistency, discriminability score, and capability alignment bias. Empirical evaluation across 15 benchmarks and 11 LLMs demonstrates that Benchmark² not only reveals substantial quality disparities among existing benchmarks but also enables the construction of streamlined test sets that maintain high evaluative performance while significantly reducing assessment scale.
Large language model (LLM) pretraining faces significant challenges, including prohibitive computational costs, opaque scaling laws, and a lack of practical guidance for large-scale distributed training. Method: This work systematically investigates the performance scaling mechanisms of LLM pretraining pipelines at the hundred-node scale, focusing on three key directions: optimization of distributed training architecture, cross-node efficient dataset management, and deep scaling of data parallelism—aiming to maximize GPU resource utilization. Through empirical analysis, we quantify the interplay among communication overhead, I/O bottlenecks, and parallelism degree, and establish a reproducible framework for large-scale training performance tuning. Contribution/Results: Our study bridges a critical gap in the public literature on engineering practices for ultra-large-scale LLM training, delivering a practical, deployable technical pathway and actionable guidelines for efficient pretraining on thousand-GPU clusters.
Existing language model evaluation benchmarks suffer from overly abstract criteria, coarse granularity, and coverage bias. To address these limitations, we propose BiGGen Bench—the first generative evaluation benchmark targeting nine fine-grained capabilities (e.g., reasoning consistency, factual controllability) across 77 diverse tasks. Our method introduces instance-level dynamic evaluation criteria and a language model self-assessment paradigm, enabling capability disentanglement and balanced assessment via collaborative scoring by multiple evaluator LMs, task-aware prompt engineering, and a structured evaluation protocol. We further develop an extensible, reproducible, and fully open-source automated evaluation framework. Comprehensive evaluation of 103 state-of-the-art models reveals critical capability bottlenecks across dimensions. All code, data, and results are publicly released.
This work addresses a critical limitation in conventional large language model (LLM) evaluation, which treats benchmark datasets as homogeneous aggregates and overlooks the heterogeneity among samples in cognitive, linguistic, and task-related attributes. The authors propose a dataset-centric meta-evaluation framework that introduces fine-grained, sample-level annotations across five dimensions: cognitive demand, language quality, task characteristics, contextual dependency, and ethical safety. For the first time, this approach enables multidimensional auditing of widely used benchmarks such as MMLU and ARC. By allowing dynamic subset composition aligned with specific evaluation objectives, the framework uncovers the diversity obscured by aggregate accuracy metrics and establishes a composable evaluation paradigm tailored to targeted capabilities—such as reasoning depth or ethical sensitivity—thereby substantially enhancing the precision and interpretability of LLM assessments.
This work investigates the reliability of micro-benchmarking for language models: whether extremely small subsets can stably reproduce the model rankings obtained on full benchmarks. We propose a meta-evaluation metric that quantifies a micro-benchmark’s ability to correctly rank models according to their full-benchmark performance differences. Combining statistical resampling with multi-benchmark experiments (MMLU-Pro, BIG-bench Hard), we systematically compare sorting consistency across diverse subset selection strategies and random sampling. Key findings reveal that existing micro-benchmarks lack stability in distinguishing models with similar capabilities; approximately 250 samples are required to ensure robust ranking, whereas with only 25 examples, over half of pairwise comparisons among 8B-parameter models fail. This study provides the first fine-grained trade-off analysis between micro-benchmark scale and ranking consistency, offering both theoretical foundations and practical guidelines for efficient, trustworthy model evaluation.
This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.
本文通过分析14,767篇论文,探讨了大型语言模型(LLM)评估基准的设计变化及其对模型性能期望的影响。
This work addresses the limitations of traditional static benchmarks—prone to saturation, contamination, and high updating costs—and the susceptibility of existing large language model (LLM) auto-scoring methods to prompt sensitivity and bias. It proposes the first three-stage framework that evaluates LLMs’ *benchmark design capability* rather than merely their question-answering performance. The approach leverages structured domain cards for extraction, quota-based multi-model collaborative item generation, and scoring via precise, numerical, and symbolic verifiers combined with psychometric analysis. From nine domains, it generates 16.7K items (retaining 15K core items) and constructs a designer–responder matrix with 152K scoring records. Empirical results reveal only a moderate correlation between design and answering abilities (Spearman ρ ≈ 0.37) and a strong negative association between invalid items and discrimination (r ≈ −0.62), demonstrating the framework’s effectiveness for scalable, cross-modal, and multilingual benchmark auditing.