Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how efficiency-oriented reasoning training affects the faithfulness and monitorability of chain-of-thought (CoT) in large language models. To systematically evaluate the differential effects under varying reasoning length pressures, we fine-tune models using three length constraint strategies: fixed budgets, per-sample objectives, and group relative rewards. Our analysis reveals that efficiency gains primarily compromise faithfulness by degrading reasoning consistency. Nevertheless, monitorability remains robust even with short CoTs, as models continue to accurately reflect the causal impact of input interventions on outputs. This work provides critical empirical evidence for balancing reasoning efficiency with safety and controllability.
📝 Abstract
Chain-of-thought (CoT) reasoning allows humans to inspect how large language models reach their answers, and oversee model behaviour. This reasoning comes at an increased inference cost, motivating efficient methods that train models to solve tasks using fewer tokens. However, a common concern is that such training may cause models to skip important reasoning steps, so the CoT no longer faithfully reflects the model's decision. It is unclear whether or when this occurs in practice, since different efficiency methods apply length pressure to models' CoT in distinct ways, and faithfully explaining a model's decision takes more tokens on some tasks than others. To understand these dynamics, we fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward. We evaluate how efficient reasoning affects CoT faithfulness (i.e., how well the CoT reflects model decisions on related inputs) and monitorability (i.e., whether the CoT reveals when input interventions alter the output). We find that it affects faithfulness and monitorability differently. Faithfulness falls in most settings, primarily because the trained models are less consistent. Monitorability is more robust, as models keep acknowledging the influence on their answer even when the CoT is much shorter.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought
Efficient Reasoning
Faithfulness
Monitorability
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought Faithfulness
Monitorability
Efficient Reasoning Training
Length Pressure
Fine-tuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Samuel Lewis-Lim
School of Computer Science, University of Sheffield
Xingwei Tan
Xingwei Tan
Research Associate
Natural Language Processing
M
Mario Sanger
AstraZeneca
Zhixue Zhao
Zhixue Zhao
University of Sheffield
Model EditingAI SafetyInterpretabilityModel Compression
N
Nikolaos Aletras
School of Computer Science, University of Sheffield