LLM Benchmarking via Representation Multi-task Learning

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of insufficient cross-domain information utilization and the difficulty of defining composite scores in large language model evaluation. We propose a statistical framework integrating representation-based multitask learning with item response theory. By designing computationally efficient and minimax optimal estimators, our approach enables effective cross-domain information sharing, establishing a rigorous measurement foundation for both domain-specific and overall capability assessment. Combined with asymptotic statistical inference techniques, experiments on simulated and real-world datasets demonstrate that the proposed method significantly outperforms existing baselines. Furthermore, it effectively reveals domain heterogeneity and strong cross-domain dependencies, offering a novel paradigm for the multidimensional capability evaluation of large language models.
📝 Abstract
Quantifying and evaluating the capabilities of Large Language Models (LLMs) remains a fundamental challenge in modern data science and artificial intelligence. In this paper, we consider LLM evaluation based on their performance across items in multiple benchmark domains (e.g., mathematical reasoning and coding) within a leaderboard framework. Our goal is to address two core questions: (1) How do we derive more accurate domain-specific scores by borrowing information across domains? and (2) How do we define and estimate an overall score that aggregates performance across multiple domains? To solve these problems, we propose a novel statistical framework based on representation Multi-task Learning (MTL) and an item response theory model. Specifically, we define overall and domain-specific LLM traits through an Item Response Theory (IRT) model, and propose an MTL approach to estimate these traits from item-level response data. We develop a computationally efficient estimator and establish its minimax optimality under certain asymptotic regimes. This framework provides a rigorous measurement foundation for systematic LLM evaluation. We conduct extensive simulations, demonstrating the superior performance of the proposed method over competing methods. Crucially for the Applications and Case Studies section, we apply the proposed framework to MMLU response data from the Hugging Face Open LLM Leaderboard, covering 4,272 LLMs and 13,232 items across 56 subjects. The empirical analysis reveals substantial heterogeneity in domain size and difficulty, together with strong positive cross-domain dependence, highlighting the practical value and substantive insights generated by our approach.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
LLM Benchmarking
Multi-task Learning
Item Response Theory
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Representation Multi-task Learning
Item Response Theory
LLM Benchmarking
Minimax Optimality
Cross-domain Estimation
🔎 Similar Papers
No similar papers found.