🤖 AI Summary
This study investigates the representation of consonant cluster reduction—a prevalent feature in African American English—within mainstream speech models and its implications for automatic speech recognition fairness. Employing speaker-independent probing across model layers, the authors design tasks to detect surface consonant clusters and recover underlying segments, comparing self-supervised (wav2vec 2.0) and supervised (Whisper-small) architectures. The findings reveal, for the first time, that such reduction is not merely phonetic deletion but is encoded as a structured, gradient phonological variation within the models: they accurately distinguish reduced from canonical forms, and crucially, retain sub-phonemic cues to underlying stops even when surface segments are absent. This demonstrates that contemporary speech models possess a nuanced, structured phonological representation of non-mainstream dialectal patterns.
📝 Abstract
Self-supervised and supervised speech models are increasingly used to investigate which linguistic information their internal representations encode, and at what level of abstraction they encode it. One underexplored phenomenon is consonant cluster reduction (CCR) in African American English (AAE), a widespread phonological process and a source of automatic speech recognition (ASR) disparity. To examine how CCR is represented, we conduct speaker-independent layer-wise probing of wav2vec2-base and Whisper-small using two tasks: segmental reduction detection and segmental restoration of underlying cluster identity. Both models distinguish reduced and canonical forms with high accuracy. Crucially, reduced segments retain cues to their underlying stops, indicating that CCR is encoded as structured gradient phonological variation rather than simple segmental deletion. These results demonstrate structured phonological encoding of AAE CCR patterns in modern speech models.