CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether cognitive ability scores of large language models form a stable, generalizable multidimensional structure. To this end, we introduce CogArena, a benchmark integrating 13 procedurally generated cognitive tasks, and develop the first multimethod validation framework combining behavioral trait analysis, covariance structure modeling, cross-intervention experiments with frozen models, and cross-model-family prediction. Our findings reveal a common principal axis underlying model cognitive performance, yet the hypothesized five-dimensional structure proves unstable. Theoretically aligned prompting yields only marginal diagonal advantages and fails to pass rigorous statistical tests or generalize to novel model families. This work establishes a systematic validation paradigm and empirical foundation for evaluating cognitive capabilities in language models.
📝 Abstract
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.
Problem

Research questions and friction points this paper is trying to address.

cognitive ability
large language models
dimensional structure
generalization
evaluation benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimethod evaluation
cognitive architecture
large language models
procedural benchmarking
dimensional validity
D
Dengzhe Hou
Graduate School of Information Sciences, Tohoku University; Unprecedented-scale Data Analytics Center, Tohoku University
L
Lingyu Jiang
Graduate School of Information Sciences, Tohoku University
Fangzhou Lin
Fangzhou Lin
Texas A&M University, Worcester Polytechnic Institute, Tohoku University
LLM/VLMcomputer visionpoint cloud
K
Kazunori D Yamada
Graduate School of Information Sciences, Tohoku University; Unprecedented-scale Data Analytics Center, Tohoku University