🤖 AI Summary
This study investigates whether cognitive ability scores of large language models form a stable, generalizable multidimensional structure. To this end, we introduce CogArena, a benchmark integrating 13 procedurally generated cognitive tasks, and develop the first multimethod validation framework combining behavioral trait analysis, covariance structure modeling, cross-intervention experiments with frozen models, and cross-model-family prediction. Our findings reveal a common principal axis underlying model cognitive performance, yet the hypothesized five-dimensional structure proves unstable. Theoretically aligned prompting yields only marginal diagonal advantages and fails to pass rigorous statistical tests or generalize to novel model families. This work establishes a systematic validation paradigm and empirical foundation for evaluating cognitive capabilities in language models.
📝 Abstract
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.