Rethinking Reasoning Paths as Phase-Structured Trajectories

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing methods in disentangling reasoning path quality from question-level noise within hidden states, which introduces analytical bias. To overcome this, we propose PAIR, a framework that models reasoning as phased trajectories within a fixed question. PAIR introduces a novel relative phase alignment mechanism that contrasts successful and failed trajectories sharing identical questions and phases, thereby effectively isolating path-quality signals. By integrating normalized progress mapping, phase-specific probe training, and latent-space directional steering, the proposed method substantially improves trajectory ranking for identical questions and Best-of-N selection performance. Furthermore, its causal efficacy is validated through phase-guided interventions.
📝 Abstract
Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.
Problem

Research questions and friction points this paper is trying to address.

reasoning paths
hidden states
large language models
path quality
correctness probes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phase-Structured Trajectories
Intra-question Reasoning
Path-Quality Probes
Trajectory Alignment
Causal Steering
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhenghao He
Department of Computer Science, University of Virginia
Guangzhi Xiong
Guangzhi Xiong
University of Virginia
Sanchit Sinha
Sanchit Sinha
University of Virginia
Natural Language ProcessingMachine LearningComputer Vision
B
Bohan Liu
Department of Computer Science, University of Virginia
Wenqian Ye
Wenqian Ye
University of Virginia
Machine LearningAlignmentAgentic AIEmbodied Intelligence
A
Aidong Zhang
Department of Computer Science, University of Virginia