🤖 AI Summary
This study addresses the inherent trade-off between mitigating cognitive biases and preserving reasoning capabilities in large language models (LLMs) by proposing DIY, a novel debiasing framework. For the first time, this work systematically transfers five human cognitive interventions validated in social psychology to LLMs, establishing a cognitive grounding mechanism. These interventions are implemented through three paradigms: demonstration via in-context exemplars, training through instruction tuning, and correction using guided self-reflection. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across multiple benchmarks, reducing average bias to merely 2% and decreasing bias on unseen dimensions by 14.8%, while maintaining 90% reasoning accuracy. Consequently, this framework effectively optimizes the balance between debiasing efficacy and reasoning proficiency.
📝 Abstract
Bias has long been studied in social psychology and cognitive science, where decades of research have produced a body of validated interventions that reduce stereotypical thinking and prejudiced responses in humans. We propose Debias It Yourself (DIY), a cognitively grounded framework that translates five such interventions into debiasing procedures for large language models and delivers them through three established paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, eleven debiasing baselines, and three reasoning benchmarks, Train+Revise and Revise alone attain the top two average ranks, lead the bias-reasoning tradeoff (mean bias as low as 2% at 90% reasoning accuracy), and reduce bias on unseen dimensions by up to 14.8%. Our code and data are publicly available.