🤖 AI Summary
This study investigates whether language models exhibit a human-like general intelligence factor and questions the validity of the prevailing development paradigm that targets a single, unified capability. To address this, we conduct the first large-scale direct test of the general intelligence hypothesis at this scale, performing factor analysis on the performance of 1,618 models across 456 benchmarks. By employing latent variable modeling, missing value imputation, and multi-source data triangulation to effectively handle sparse data, we explore the interpretable structure of machine intelligence. Our findings reveal that a general intelligence factor accounts for only approximately 70% of the performance variance and lacks a unified semantic theme. This empirically demonstrates that strategies aimed at enhancing a singular general capability are unfounded, thereby challenging the core assumptions underlying current large model development.
📝 Abstract
A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The $g$ factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.