How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Processing

📅 2026-02-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how large language models reject incorrect answers in factual queries. By analyzing the evolution of hidden-state trajectories for correct versus incorrect completions, the authors identify a two-stage geometric pattern: mid-layer directional rotation separates representations, followed by asymmetric commitment in later layers—termed “rotational divergence followed by late asymmetric commitment.” This pattern serves as an observable signature of factual constraint processing and supports an explanatory framework based on distributed trajectory dynamics rather than localized circuit mechanisms. Through hidden-state analysis, logit-lens probing, single-layer activation patching, and cross-model validation, the mechanism is consistently reproduced across six decoder-only Transformers ranging from 1B to 13B parameters, indicating that factual judgment relies on geometric structures accumulated across multiple layers and refuting hypotheses positing recall via single-layer local circuits.
📝 Abstract
When a language model is fed a wrong answer, what happens inside the network? Current understanding treats truthfulness as a static property of individual-layer representations-a direction to be probed, a feature to be extracted. Less is known about the dynamics: how internal representations diverge across the full depth of the network when the model processes correct versus incorrect continuations. We introduce forced-completion probing, a method that presents identical queries with known correct and incorrect single-token continuations and tracks five geometric measurements across every layer of four decoder-only models(1.5B-13B parameters). We report three findings. First, correct and incorrect paths diverge through rotation, not rescaling: displacement vectors maintain near-identical magnitudes while their angular separation increases, meaning factual selection is encoded in direction on an approximate hypersphere. Second, the model does not passively fail on incorrect input-it actively suppresses the correct answer, driving internal probability away from the right token. Third, both phenomena are entirely absent below a parameter threshold and emerge at 1.6B, suggesting a phase transition in factual processing capability. These results show that factual constraint processing has a specific geometric character-rotational, not scalar; active, not passive-that is invisible to methods based on single-layer probes or magnitude comparisons.
Problem

Research questions and friction points this paper is trying to address.

Transformers
factual constraint processing
rotational dynamics
hidden-state geometry
wrong answer rejection
Innovation

Methods, ideas, or system contributions that make the work stand out.

rotational dynamics
factual constraint processing
hidden-state geometry
decoder-only transformers
distributed representation
🔎 Similar Papers