When Entanglement Lower-Bounds Disparity: Auditing and Repairing Demographic Fairness in Audio Understanding Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入TRIAD审计网格和ORCA适配器,解决了语音识别技术对不同人口统计学特征的不公平问题,显著降低了偏差。
📝 Abstract
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage $Λ$ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks $Λ$ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
Problem

Research questions and friction points this paper is trying to address.

Demographic Fairness
Audio Understanding Models
Speech Recognition
Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

TRIAD
voice-semantic leakage
axis-fidelity functionals
ORCA
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.