Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between mitigating cognitive biases and preserving reasoning capabilities in large language models (LLMs) by proposing DIY, a novel debiasing framework. For the first time, this work systematically transfers five human cognitive interventions validated in social psychology to LLMs, establishing a cognitive grounding mechanism. These interventions are implemented through three paradigms: demonstration via in-context exemplars, training through instruction tuning, and correction using guided self-reflection. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across multiple benchmarks, reducing average bias to merely 2% and decreasing bias on unseen dimensions by 14.8%, while maintaining 90% reasoning accuracy. Consequently, this framework effectively optimizes the balance between debiasing efficacy and reasoning proficiency.
📝 Abstract
Bias has long been studied in social psychology and cognitive science, where decades of research have produced a body of validated interventions that reduce stereotypical thinking and prejudiced responses in humans. We propose Debias It Yourself (DIY), a cognitively grounded framework that translates five such interventions into debiasing procedures for large language models and delivers them through three established paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, eleven debiasing baselines, and three reasoning benchmarks, Train+Revise and Revise alone attain the top two average ranks, lead the bias-reasoning tradeoff (mean bias as low as 2% at 90% reasoning accuracy), and reduce bias on unseen dimensions by up to 14.8%. Our code and data are publicly available.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Cognitive Bias Mitigation
Debiasing
Stereotypical Thinking
Bias-Reasoning Tradeoff
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cognitive Bias Mitigation
Large Language Models
Instruction Tuning
Guided Self-Revision
Debiasing Framework
🔎 Similar Papers
No similar papers found.