🤖 AI Summary
Current evaluations of large language models predominantly focus on isolated tasks, lacking a systematic capability framework, which hinders cross-study comparability and leaves critical coverage gaps. This work proposes the first three-tier capability taxonomy grounded in human cognitive science, encompassing 14 capability domains and 91 sub-skills, thereby shifting the unit of analysis from tasks to structured competencies. Leveraging multi-model collaborative annotation, consensus mechanisms, and arbitration protocols, the study systematically maps and statistically analyzes nearly 16,000 papers published between 2023 and 2025. The findings reveal a pronounced concentration of research on language and reasoning, with over 60% of capability domains accounting for less than 2% of studies. Additionally, the analysis identifies recurrent co-occurring capability clusters and quantifies their association strengths, offering empirically testable hypotheses for model evaluation, training, and transfer.
📝 Abstract
Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and makes coverage gaps difficult to identify.
We introduce a multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers. Human cognitive science guides capability definition and organization, not LLM architecture. Layer assignments draw on developmental precedence and hypothesized functional support, while human-origin constructs are adapted to observable model behavior.
To demonstrate operational utility, we screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS between 2023 and 2025 and mapped 15,934 LLM-focused papers through multi-model annotation, consensus, and arbitration. Direct research attention concentrated on Language-Semantic Competence (3,551; 22.3%), Reasoning (3,388; 21.3%), Planning and Decision-Making (2,149; 13.5%), and Perception (1,954; 12.3%), whereas six domains appeared in fewer than 2% of papers. Within domains, the most frequent subskill had a median prevalence of 97.9% and appeared in at least 90% of papers in 10 of 14 domains. Language-Semantic Competence and Reasoning formed the highest-volume pair (n = 1,864; 11.7%; lift = 2.47), whereas Theory of Mind and Social Reasoning and Interaction showed the highest lift among pairs with at least 20 co-occurrences (n = 62; lift = 30.84).
By shifting the unit of analysis from isolated tasks to structured capabilities, the taxonomy supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.