🤖 AI Summary
This work identifies a fundamental ceiling in the knowledge-intensive capabilities of decoder-only autoregressive language models: despite scaling parameters from 70M to 30B and consistent reduction in cross-entropy loss, accuracy on knowledge retrieval and the MMLU mathematics subset remains stagnant at 19–20%—well below random chance—contrasting sharply with standard scaling behavior on procedural tasks. We systematically evaluate OPT and Pythia families (70M–30B) using multidimensional metrics and attention pattern swapping experiments. Our analysis provides the first empirical evidence of severe decoupling between training loss and knowledge utilization capacity, and demonstrates pronounced diminishing returns to parameter scaling for such tasks. Furthermore, we show that current decoder-only architectures exhibit high sensitivity to attention mechanisms, suggesting that the bottleneck in knowledge integration may be intrinsic to the decoder architecture itself.
📝 Abstract
We document empirical capability ceilings in decoder-only autoregressive language models across knowledge-intensive tasks. Systematic evaluation of OPT and Pythia model families (70M-30B parameters, spanning 240 times scaling) reveals that knowledge retrieval tasks show negligible accuracy improvement despite smooth loss reduction. On MMLU mathematics benchmarks, accuracy remains flat at 19-20% (below 25% random chance) across all scales while cross-entropy loss decreases by 31%. In contrast, procedural tasks like arithmetic show conventional scaling where both metrics improve together. Attention intervention experiments reveal high sensitivity to perturbation: swapping attention patterns between models causes catastrophic performance collapse (complete accuracy loss) rather than graceful degradation. These measurements have immediate engineering implications: for knowledge-intensive applications using OPT and Pythia architectures, parameter scaling beyond 1-2B offers minimal accuracy gains despite continued loss improvement. Our findings quantify capability-specific scaling failures in these model families to inform resource allocation decisions. Whether these patterns reflect fundamental constraints of decoder-only architectures or implementation-specific limitations remains an open question requiring investigation across diverse architectural approaches.