Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the contradictory findings regarding the impact of chain-of-thought reasoning on counterfactual fairness in reasoning language models. Focusing on models such as QwQ, we conduct ablation experiments to investigate fairness disparities between thinking and non-thinking modes in high-stakes decision-making tasks. Methodologically, we propose the Counterfactual Deep Probability Gap (CDPG) and the Bias Transition Matrix (BTM) to quantitatively trace the evolutionary trajectories and state transition mechanisms of bias throughout the reasoning process. Our findings reveal that while reasoning can mitigate certain pre-existing biases, it simultaneously introduces approximately five times more novel biases, which further propagate and amplify with increasing reasoning depth. These results uncover an asymmetric dual effect of reasoning on model fairness.
📝 Abstract
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinking baseline and creates new flips at near-saturating model confidence. In all nine (model, dataset) combinations, the created flips outnumber the resolved flips by roughly 5 times. To explain the effect, we treat the thinking trace itself as a measurable site of fairness change and study it through two dynamic instruments: 1) We propose Counterfactual Depth Probability Gap (CDPG) to track bias evolution along thinking depth, and observe that bias propagates and amplifies with thinking. 2) We also formulate the Bias Transition Matrix (BTM) to show how predictions of counterfactual pairs change from non-thinking to thinking, and find that the asymmetric dual effect originates in the pair-state joint transition.
Problem

Research questions and friction points this paper is trying to address.

reasoning language models
counterfactual fairness
bias amplification
thinking tokens
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reasoning Language Models
Counterfactual Fairness
Thinking Trace Analysis
Counterfactual Depth Probability Gap (CDPG)
Bias Transition Matrix (BTM)
🔎 Similar Papers
💼 Related Jobs
No related jobs found.