Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling

πŸ“… 2026-05-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study reveals a phase transition in the relationship between reasoning ability and truthfulness during language model scaling: below a critical scale of approximately 3.5 billion parameters, the two exhibit a negative correlation, which reverses to a positive correlation beyond this thresholdβ€”a phenomenon invisible to conventional loss curves. Analyzing 63 base models through cross-architecture benchmark scoring, sparse regression-based ODE modeling, attention head statistics, and activation interventions, the work identifies and quantifies this transition for the first time. The authors demonstrate that width normalization eliminates the negative correlation, while data refinement, knowledge distillation, and architectural innovations substantially advance the onset of the cooperative regime. An open-source diagnostic and prediction platform is released, requiring no internal model access.
πŸ“ Abstract
Scaling laws predict loss from compute but not how capabilities interact. We measure the coupling between reasoning and truthfulness across 63 base models from 16 families and find a regime change invisible to loss curves: below a family-dependent critical scale $N_c$, capabilities anticorrelate; above it, they cooperate. $N_c \approx 3.5$B parameters [2.9B, 13.4B] (bootstrap 95% CI), but model size is not the only variable that determines phase. Architecture, data curation, and training recipe each shift $N_c$ independently: curated training eliminated the coupling dip between Qwen generations ($0.025 \to 0.830$ at matched scale), Gemma-4 at 4B achieves coupling 0.871, characteristic of 13B+ standard-trained models, through distillation and architectural innovation, and Phi at 1B matches web-trained coupling at 10B through data curation alone. Width normalization eliminates the anticorrelation across all tested families, supporting an output-projection bottleneck. Internally, 38 of 40 models show zero competing attention heads. A sparse-regression ODE cross-predicts held-out Llama-2 at 5.6% error. The diagnostic requires no model internals -- only public benchmark scores across a model family. The cooperative regime extends to the frontier ($r = +0.72$, 34 models, 10 labs). Code, data, and an open-source activation-steering tool for any open-weight model are released alongside an interactive dashboard that diagnoses any model's coupling phase, suggests concrete interventions (data curation, width, benchmark rotation), and provides ODE scaling predictions, frontier diagnostics, and eigenstructure analysis: https://zehenlabs.com/cape/.
Problem

Research questions and friction points this paper is trying to address.

scaling laws
truthfulness
reasoning
phase transition
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

alignment phase transition
truthfulness-reasoning coupling
scaling laws
width normalization
model diagnostics
πŸ”Ž Similar Papers