Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing LLM debiasing methods, which rely on explicit examples or fixed substitutions and struggle to capture cross-contextual biases or model the statistical dependence between outputs and biases. We propose Acmite, a lightweight concept-guided framework that structures stereotypes into semantic concepts, filters them via Maximum Marginal Relevance, and achieves targeted debiasing through mutual information minimization. Furthermore, it introduces a conditional triggering mechanism that activates LoRA adapters only upon bias detection, effectively balancing debiasing performance with the preservation of general capabilities. Experimental results demonstrate that Acmite significantly mitigates gender bias on benchmarks such as BBQ while maintaining competitive performance on downstream tasks including ARC and GSM8K.
📝 Abstract
Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at https://anonymous.4open.science/r/Acmite-18E2/.
Problem

Research questions and friction points this paper is trying to address.

Gender Bias
Large Language Models
Debiasing
Stereotype Concepts
Statistical Dependence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gender Bias Mitigation
Mutual Information Minimization
Concept-Guided Debiasing
Maximal Marginal Relevance
LoRA Adapter
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.