Score
Implements and evaluates procedures to adapt pretrained self‑supervised models into task‑specific systems; for SSL encoders — including audio/speech models such as WavLM — this covers choosing which layers to freeze or progressively unfreeze, adding and training classifier heads, and tuning optimization and regularization schedules. Also develops and tests techniques to preserve or improve representation robustness when labeled data are scarce, for example via selective fine‑tuning schedules, augmentation, and label‑efficient training strategies.
To address the limited cross-lingual transferability and severe domain mismatch of self-supervised learning (SSL) pre-trained models in automatic speech recognition (ASR) for low-resource languages, this paper proposes a lightweight adapter method with *intermediate warm-start*. Under frozen SSL backbone constraints, only 1–5% of parameters are fine-tuned. A two-stage progressive adaptation jointly optimizes adapter architecture and downstream model initialization. The novel intermediate warm-start mechanism mitigates speech feature distribution shift, substantially improving generalization to unseen languages. Evaluated on the ML-SUPERB benchmark, our approach achieves up to 28% relative reduction in character/phone error rates over standard efficient fine-tuning, significantly alleviating the bottleneck in low-resource cross-lingual ASR adaptation.
This study investigates whether the self-supervised speech representation model HuBERT can directly support speech inpainting—reconstructing missing or corrupted speech segments—without task-specific fine-tuning. We propose an encoder–decoder collaborative modeling framework: a frozen or fine-tuned HuBERT encoder is coupled with a HiFi-GAN vocoder decoder to jointly model context-aware waveform generation. Our key contribution is the first explicit alignment of self-supervised pretraining objectives with speech inpainting, supporting both known and unknown mask locations, as well as single- and multi-speaker scenarios. Experiments demonstrate that fine-tuning HuBERT achieves precise reconstruction of up to 400-ms segments in single-speaker settings; in multi-speaker settings, freezing HuBERT while optimizing HiFi-GAN significantly improves naturalness and intelligibility. Both objective metrics (e.g., PESQ, STOI) and subjective listening evaluations confirm the effectiveness of our approach.
It remains unclear whether speech-based self-supervised pre-trained models can effectively transfer to bioacoustic tasks—challenging the prevailing assumption that domain-specific pretraining on animal vocalizations is necessary. Method: We systematically evaluate speech models (Wav2Vec 2.0, HuBERT) against animal-call-specific models (e.g., AVES) across diverse bioacoustic datasets (FreeSound, BirdVox) on species identification and sound event detection, with and without ASR fine-tuning. Contribution/Results: Speech models match or surpass animal-call-specific models on most tasks; ASR fine-tuning yields marginal gains, indicating that general acoustic representations already possess strong bioacoustic adaptability. This work provides the first empirical evidence refuting the necessity of domain-specific pretraining for bioacoustics, establishing a “light-fine-tuning, high-efficiency” paradigm. It offers a scalable, resource-efficient methodology for low-data animal acoustic modeling, advancing transfer learning in ecological audio analysis.
To address the dual challenges of “new-language enhancement” and “original-language capability preservation” in cross-lingual adaptation of self-supervised speech models, this paper proposes a LoRA-driven framework for language-incremental expansion. Methodologically, it integrates Low-Rank Adaptation (LoRA) with a dual-track capability retention strategy: (i) multilingual data mixing during fine-tuning and (ii) k-means re-clustering to optimize the discrete representation space. The approach is instantiated on the mHuBERT architecture to enable efficient Chinese extension. Experiments demonstrate that only 0.3% of mHuBERT’s parameters require tuning for Chinese integration, yielding a MOS improvement of 1.6 and a relative WER reduction of 61.72%, while preserving zero performance degradation across all pre-existing language tasks. To our knowledge, this is the first work to systematically introduce LoRA into progressive multilingual expansion of self-supervised speech models, achieving a favorable trade-off among parameter efficiency, cross-lingual compatibility, and capability stability.
To address the inefficiency and overfitting issues of large-scale self-supervised speech models (e.g., wav2vec 2.0) under edge-device resource constraints—particularly in multilingual and multi-task scenarios—this paper proposes S³-Router, a novel dynamic sparse routing framework that abandons conventional weight fine-tuning and instead optimizes only the inter-layer connection topology. Theoretically and empirically, we demonstrate for the first time that pruning ≤10% of connections yields superior downstream performance compared to full-parameter fine-tuning. S³-Router unifies several critical capabilities: efficient model adaptation, joint multilingual/multi-task modeling, ASR model pruning, and representation interpretability analysis. On low-resource ASR tasks, it achieves significant accuracy gains while drastically reducing inference FLOPs and memory footprint. The method is inherently deployment-friendly on edge devices and exhibits strong generalization and robustness across diverse domains and languages.
This study investigates the language sensitivity of neural audio codecs (NACs) and self-supervised learning (SSL) speech models in multilingual settings, addressing whether separate models must be trained for each language. By fixing the pretraining language of either the NAC or SSL model and systematically evaluating downstream task performance, the work reveals—for the first time—that the NAC’s training language has negligible impact on performance, whereas alignment between the SSL pretraining language and the target language is critical. These findings demonstrate that a single NAC can be effectively reused across languages, substantially reducing the training cost of multilingual SSL systems without compromising performance. This insight establishes a new paradigm for efficiently building multilingual speech models.
This work addresses the vulnerability of self-supervised speech representations to positional embedding interference during fine-tuning for speech enhancement, which often leads models to over-rely on positional cues rather than actual speech content. To mitigate this issue, the authors propose a position-invariant fine-tuning strategy that integrates speed perturbation with zero-padding and introduces a soft-DTW alignment loss to effectively decouple content from positional information. The proposed approach significantly improves speech enhancement performance under noisy conditions, accelerates model convergence, and yields superior results on downstream tasks, thereby demonstrating the effectiveness and practicality of position-invariant fine-tuning in leveraging self-supervised speech representations.
Although self-supervised speech pretraining relies on large-scale data, effective data selection strategies remain unclear. This work systematically evaluates the impact of different pretraining subsets on automatic speech recognition (ASR) performance and proposes several filtering strategies—including random sampling, diversity-based selection, and prioritizing longer utterances—for comparative analysis. The experiments reveal that utterance length is more decisive than data diversity or total volume: using only the top 50% longest utterances surpasses the performance achieved with the full dataset while reducing pretraining time by 24%. These findings highlight the critical role of speech duration in self-supervised learning and offer a new direction for efficient speech model training.
This work addresses the degradation of semantic information in self-supervised speech representations under noisy conditions, where existing adaptation modules often preserve acoustic details at the expense of linguistic content during joint training. To mitigate this issue, the authors propose a decoupled semantic aggregation strategy grounded in phoneme mutual information. Specifically, a pre-trained and frozen language aggregation layer is employed to explicitly maximize the mutual information between learned representations and phoneme labels, thereby effectively preserving linguistic content during speech enhancement. Integrating information-theoretic measures, a dynamic aggregation mechanism, and a decoupled training framework, the proposed method significantly reduces word error rate (WER) and outperforms end-to-end jointly optimized baselines.
This study challenges the prevailing assumption that performance gains from supervised fine-tuning (SFT) in downstream tasks of speech foundation models stem primarily from methodological improvements. Instead, it systematically evaluates eight SFT variants across nine pretrained checkpoints of wav2vec 2.0, HuBERT, and WavLM on three SUPERB classification tasks, incorporating multiple random seeds to assess stability and transferability. The findings reveal that SFT’s apparent advantages are highly contingent on specific pretrained instances and random seeds, with optimal configurations showing little consistency or generalizability across checkpoints. These results suggest that most reported gains arise from favorable instance–seed matching rather than genuine improvements in model capacity or upper-bound performance, thereby questioning the universality of SFT enhancements.