🤖 AI Summary
This study addresses the significant performance disparities of automatic speech recognition (ASR) systems across different accents, a phenomenon whose underlying mechanisms remain poorly understood. Focusing on the Wav2Vec2 model, the authors identify—for the first time—an 8-dimensional low-dimensional subspace in the third transformer layer that densely encodes accent-related information and exhibits a notable correlation with word error rate (r = 0.26). Through representational analysis, subspace projection, controlled perturbations, and linear attenuation experiments, they demonstrate that targeted perturbations within this subspace exacerbate performance degradation (r = 0.32). Surprisingly, simple attenuation of the subspace slightly worsens fairness, challenging the prevailing assumption that erasing accent features inherently improves ASR equity.
📝 Abstract
ASR systems exhibit persistent performance disparities across accents, yet the internal mechanisms underlying these gaps remain poorly understood. We introduce ACES, a representation-centric audit that extracts accent-discriminative subspaces and uses them to probe model fragility and disparity. Analyzing Wav2Vec2-base with five English accents, we find that accent information concentrates in a low-dimensional early-layer subspace (layer 3, k=8). Projection magnitude correlates with per-utterance WER (r=0.26), and crucially, subspace-constrained perturbations yield stronger coupling between representation shift and degradation (r=0.32) than random-subspace controls (r=0.15). Finally, linear attenuation of this subspace however does not reduce disparity and slightly worsens it. Our findings suggest that accent-relevant features are deeply entangled with recognition-critical cues, positioning accent subspaces as vital diagnostic tools rather than simple "erasure" levers for fairness.