🤖 AI Summary
Existing definitions of Artificial General Intelligence (AGI) rely on arithmetic means, implicitly assuming compensatory capability across domains—leading to overestimation of specialized models and neglecting domain dependence and capability imbalance. Method: We propose a consistency-based AGI evaluation framework grounded in the Cattell–Horn–Carroll (CHC) cognitive architecture. It constructs an Area-Under-Curve (AUC) metric via continuous integration of generalized means—from arithmetic to harmonic—to quantify cross-domain capability balance and synergistic robustness, explicitly modeling non-compensability among domains. Contribution/Results: Unlike single-mean metrics, our approach identifies critical capability bottlenecks. Empirical evaluation on GPT-4 and GPT-5 reveals that although their arithmetic mean scores reach 24%, their AUC consistency scores are markedly lower—demonstrating a fundamental gap from true general intelligence. This work establishes the first theoretically rigorous yet empirically tractable metric for assessing capability balance in AGI evaluation.
📝 Abstract
Recent work by citet{hendrycks2025agidefinition} formalized extit{Artificial General Intelligence} (AGI) as the arithmetic mean of proficiencies across cognitive domains derived from the Cattell--Horn--Carroll (CHC) model of human cognition. While elegant, this definition assumes extit{compensability} -- that exceptional ability in some domains can offset failure in others. True general intelligence, however, should reflect extit{coherent sufficiency}: balanced competence across all essential domains. We propose a coherence-aware measure of AGI based on the integral of generalized means over a continuum of compensability exponents. This formulation spans arithmetic, geometric, and harmonic regimes, and the resulting extit{area under the curve} (AUC) quantifies robustness under varying compensability assumptions. Unlike the arithmetic mean, which rewards specialization, the AUC penalizes imbalance and captures inter-domain dependency. Applied to published CHC-based domain scores for GPT-4 and GPT-5, the coherence-adjusted AUC reveals that both systems remain far from general competence despite high arithmetic scores (e.g., GPT-5 at~24%). Integrating the generalized mean thus yields a principled, interpretable, and stricter foundation for measuring genuine progress toward AGI.