Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks

📅 2025-10-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work identifies a fundamental ceiling in the knowledge-intensive capabilities of decoder-only autoregressive language models: despite scaling parameters from 70M to 30B and consistent reduction in cross-entropy loss, accuracy on knowledge retrieval and the MMLU mathematics subset remains stagnant at 19–20%—well below random chance—contrasting sharply with standard scaling behavior on procedural tasks. We systematically evaluate OPT and Pythia families (70M–30B) using multidimensional metrics and attention pattern swapping experiments. Our analysis provides the first empirical evidence of severe decoupling between training loss and knowledge utilization capacity, and demonstrates pronounced diminishing returns to parameter scaling for such tasks. Furthermore, we show that current decoder-only architectures exhibit high sensitivity to attention mechanisms, suggesting that the bottleneck in knowledge integration may be intrinsic to the decoder architecture itself.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsNatural Language Processing: (Large) Language ModelsSearch and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
We document empirical capability ceilings in decoder-only autoregressive language models across knowledge-intensive tasks. Systematic evaluation of OPT and Pythia model families (70M-30B parameters, spanning 240 times scaling) reveals that knowledge retrieval tasks show negligible accuracy improvement despite smooth loss reduction. On MMLU mathematics benchmarks, accuracy remains flat at 19-20% (below 25% random chance) across all scales while cross-entropy loss decreases by 31%. In contrast, procedural tasks like arithmetic show conventional scaling where both metrics improve together. Attention intervention experiments reveal high sensitivity to perturbation: swapping attention patterns between models causes catastrophic performance collapse (complete accuracy loss) rather than graceful degradation. These measurements have immediate engineering implications: for knowledge-intensive applications using OPT and Pythia architectures, parameter scaling beyond 1-2B offers minimal accuracy gains despite continued loss improvement. Our findings quantify capability-specific scaling failures in these model families to inform resource allocation decisions. Whether these patterns reflect fundamental constraints of decoder-only architectures or implementation-specific limitations remains an open question requiring investigation across diverse architectural approaches.
Problem

Research questions and friction points this paper is trying to address.

Autoregressive models show negligible accuracy gains in knowledge retrieval tasks
Mathematical reasoning accuracy remains flat despite significant loss reduction
Attention patterns are highly sensitive to perturbations causing catastrophic failures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Identified negligible accuracy gains despite loss reduction
Revealed attention pattern swapping causes catastrophic performance collapse
Quantified capability-specific scaling failures to inform resource allocation
🔎 Similar Papers
No similar papers found.