Temporal Taxation Compounds Under Post-Training Compression of Whisper Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how model compression techniques, such as quantization and pruning, exacerbate the “temporal tax” imposed on marginalized groups by speech recognition systems—a deployment risk that traditional full-precision auditing fails to capture. To systematically evaluate fairness degradation in compressed Whisper models, this work pioneers the quantification of this metric by integrating Wanda pruning, INT4 HQQ quantization, knowledge distillation, and multi-dataset analysis. The findings reveal that 50% pruning doubles correction latency, while quantization causes abrupt transcription performance drops for specific accents; conversely, knowledge distillation significantly narrows inter-group error disparities across most scenarios. By exposing the latent harms of post-training compression on vulnerable populations and demonstrating the limitations of single-snapshot auditing, this research validates the distinct fairness advantages of distillation-based compression strategies.
📝 Abstract
Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which alters model weights rather than the audio signal or its feature representation, redistributes error burden across demographic groups. Across the Whisper family on Fair-Speech, Common Voice 25, and AfriSpeech-200, 50% Wanda pruning of Whisper-large-v3 sharply widens the Black/AA-vs-Asian temporal-taxation differential on Fair-Speech: the absolute word-error-rate gap between the worst- and best-served groups more than doubles; at an assumed cost of five seconds of correction effort per transcription error this is a rise from 30 to 64 seconds of correction time per minute of speech. This +111% relative increase is invariant to the assumed per-error cost, survives an audio-quality control, and is only partly mitigated by beam-search decoding, which still leaves an +86% increase. At edge model size, INT4 HQQ quantization compounds catastrophic transcript loops on West African accents by factors of five to seven. Distillation, by contrast, narrows demographic gaps in 21 of 27 evaluated settings (teacher-student pair, precision, and dataset), with the exceptions concentrated on a single model pair. We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.
Problem

Research questions and friction points this paper is trying to address.

post-training compression
demographic fairness
automatic speech recognition
temporal taxation
Whisper models
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-training compression
temporal taxation
demographic fairness
speech recognition
quantization
🔎 Similar Papers
No similar papers found.