Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models

📅 2026-05-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of discerning whether lengthy reasoning traces generated by large language models genuinely reflect their internal reasoning processes or merely constitute redundant output. To this end, the authors propose StALT, a novel metric that, for the first time, quantifies dynamic patterns in model hidden states across both temporal and layer dimensions without requiring additional training. StALT constructs a spatiotemporal magnitude statistic by analyzing hidden state trajectories, applying inter-layer saliency weighting, and measuring state transitions between adjacent tokens. Experimental results demonstrate that StALT reliably distinguishes between correct and incorrect reasoning traces across diverse models and reasoning-intensive tasks, while remaining sensitive to variations in reasoning demands, thereby validating its effectiveness and generality as an internal reasoning probe.
📝 Abstract
Large reasoning models (LRMs) generate extended solutions, yet it remains unclear whether these traces reflect substantive internal computation or merely verbosity and overthinking. Although recent hidden-state analyses suggest that internal representations carry correctness-related signals, their coarse aggregations may obscure the token and layer structure underlying reasoning computation. We investigate hidden-state transitions across decoding steps and layers, and identify a distinct spatiotemporal pattern in LRMs: successful trajectories exhibit broad temporal dynamics with localized layer-wise concentration, while this structure is weaker in non-reasoning models and knowledge-heavy domains. We formalize this characteristic as Spatiotemporal Amplitude of Latent Transition (StALT), a training-free trajectory statistic that summarizes temporal changes between adjacent tokens weighted by within-token layer saliency. Across diverse models and benchmarks, StALT reliably separates correct from incorrect trajectories in reasoning-intensive regimes, providing a competitive label-free correctness signal alongside strong output-space and length-based baselines. Intervention analyses further show that this spatiotemporal amplitude responds systematically to manipulations that increase or reduce the demand for internal reasoning, supporting its association with latent reasoning dynamics in LRMs. These findings provide empirical evidence that LRMs exhibit measurable hidden-state dynamics and offer a practical probe for understanding internal computation beyond output-based evaluation.
Problem

Research questions and friction points this paper is trying to address.

reasoning
hidden-state dynamics
large language models
correctness
internal computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatiotemporal dynamics
hidden-state transitions
latent reasoning
StALT
large reasoning models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kotaro Furuya
Research and Development Group, Hitachi, Ltd.
T
Takahito Tanimura
Research and Development Group, Hitachi, Ltd.