Interpreting Reasoning of Large Language Models via Partial Information Decomposition

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of verbose, repetitive, and error-prone reasoning trajectories exhibited by large reasoning models on complex tasks by proposing the SLIDER framework. Grounded in information theory, this method decouples the informational components across reasoning steps and pioneers the use of partial information decomposition to construct step-wise and trajectory-level Repetitive Reasoning Indices (RRI), enabling a theoretically driven quantitative assessment of reasoning redundancy. Furthermore, by integrating the RRI metric into data filtering and fine-tuning optimization, the framework effectively eliminates inefficient reasoning paths. Experimental results demonstrate that the proposed approach improves redundancy identification accuracy by over 10%, significantly enhancing the model's reasoning efficiency while preserving task performance.
📝 Abstract
Large reasoning models (LRMs) have achieved substantial improvements in solving complex mathematical problems, but often produce lengthy, repetitive, or erroneous reasoning trajectories. In this work, we introduce a new interpretability framework, SLIDER, to evaluate the quality of the reasoning process. SLIDER leverages an emerging body of work from information theory called Partial Information Decomposition to disentangle the information about the final answer between two consecutive reasoning steps into non-negative components: unique information (in preceding steps or current step), redundant information, and synergistic information. Building on this decomposition, we propose the *Step-wise Repetitive Reasoning Index (Step-RRI)*, a theoretically grounded measure that assesses whether the answer-relevant information in the current step $S_i$ is predominantly redundant with the past steps $S_{<i}$, relative to its unique and synergistic contributions. To evaluate the effectiveness of Step-RRI in detecting repetitiveness, we apply SLIDER to the redundancy class of the PRMBench dataset where Step-RRI improves step-level redundancy identification accuracy by over $10$ points compared to embedding-similarity and information-gain baselines. Next, we define *Trajectory-RRI*, an aggregate measure of repetitiveness for an individual reasoning trajectory. To demonstrate its practical relevance, we show that average Trajectory-RRI strongly correlates with actual reasoning length across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and GPT-4.1, motivating its use as a signal for improving reasoning efficiency. Finally, we introduce *Trajectory-RRI-guided data selection for fine-tuning*, demonstrating that selecting training data based on Trajectory-RRI can improve a fine-tuned model's reasoning efficiency while largely preserving its task performance.
Problem

Research questions and friction points this paper is trying to address.

Large Reasoning Models
Reasoning Trajectories
Repetitive Reasoning
Interpretability
Reasoning Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Partial Information Decomposition
Reasoning Interpretability
Redundancy Detection
Step-wise Repetitive Reasoning Index
Data Selection
🔎 Similar Papers
No similar papers found.