Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing generalization evaluations for large language models, which frequently conflate robustness with accuracy and lack fine-grained metrics for behavioral stability under input variations. To this end, this work proposes the Stability-Aware Generalization Objective (SAGO) framework, which reconceptualizes generalization as stability. SAGO quantifies behavioral fluctuations across semantically equivalent input variants through multi-axis metrics, including generation consistency and internal activation patterns. Experimental results demonstrate significant generalization instability in mainstream models, revealing that distinct failure modes are mutually independent and cross-dataset performance rankings can invert. These findings expose the fundamental limitations of relying on single-metric accuracy and establish a novel paradigm for evaluating the reliability of large language models.
📝 Abstract
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Generalization
Stability
Evaluation
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generalization Stability
SAGO Framework
Multi-Axis Evaluation
Large Language Models
Behavioral Variability
💼 Related Jobs
No related jobs found.
N
Nagham Omar
Technion – Israel Institute of Technology
M
Mahmoud Jabarin
Technion – Israel Institute of Technology
M
Maya Rozenshtein
Technion – Israel Institute of Technology
R
Rom Himelstein
Technion – Israel Institute of Technology
Avi Mendelson
Avi Mendelson
Electrical Engineering and Computer Science, Technion,
Computer systems
A
Amit LeVi
Technion – Israel Institute of Technology