$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the excessive inference overhead of chain-of-thought large language models by revealing, for the first time, that core reasoning capabilities reside primarily in the null space of non-thinking models rather than in the dominant subspace emphasized by mainstream approaches. Accordingly, this work proposes a training-free spectral null-space swapping method that achieves efficient model fusion through null-space projection and checkpoint composition, with attention entropy providing a theoretical explanation. Experimental results demonstrate that the proposed approach reduces token consumption by an average of 27.4% while improving accuracy by 1.0%, thereby establishing a new Pareto frontier for performance and efficiency.
📝 Abstract
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.
Problem

Research questions and friction points this paper is trying to address.

reasoning efficiency
token cost
large language models
chain-of-thought
model optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Null Space
Training-free Model Composition
Reasoning Efficiency
Spectral Analysis
Chain-of-Thought
🔎 Similar Papers
No similar papers found.