🤖 AI Summary
This study addresses the reliance of existing speech self-supervised learning on complex, manually designed prediction objectives—such as contrastive learning or discretization—which result in cumbersome training pipelines. To overcome this limitation, this work proposes SIGReg, a framework that eliminates engineered objective generation mechanisms by directly predicting continuous representations at masked positions. Grounded in the Joint Embedding Predictive Architecture (JEPA), the method introduces Gaussian regularization to constrain the spatial distribution of representations, effectively preventing representation collapse and substantially simplifying the pretraining procedure. Experimental evaluations on LibriSpeech demonstrate that the proposed framework achieves a word error rate of 6.89% on automatic speech recognition tasks and a character error rate of 25.87% on slot filling, significantly outperforming existing baseline methods.
📝 Abstract
Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.