Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue whereby aggregate word error rates in pruned speech large language models obscure group-level performance disparities, presenting the first systematic investigation into how audio encoder pruning affects the fairness of SLAM-ASR. Utilizing the Fair-Speech and Common Voice datasets, we conduct multi-scale pruning experiments integrated with LoRA adaptation. Our findings reveal that pruning significantly exacerbates performance gaps across demographic groups, and that LoRA optimization may further amplify such inequalities. To mitigate these effects, this work proposes incorporating worst-group error rates into deployment decision criteria and emphasizes the necessity of monitoring evaluation metrics at the group level. Ultimately, this research uncovers the latent bias risks concealed behind efficiency improvements, advocating for fairness-aware practices when compressing speech foundation models for real-world deployment.
📝 Abstract
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group's error rate as an explicit criterion.
Problem

Research questions and friction points this paper is trying to address.

Speech-LLMs
Model Pruning
Fairness
Demographic Disparities
Word Error Rate
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech-LLM Pruning
Algorithmic Fairness
Demographic Disparities
Word Error Rate (WER)
LoRA Adaptation
🔎 Similar Papers
No similar papers found.