Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the adaptation bottleneck of conventional LoRA in low-resource speech recognition, where task-specific guidance is lacking. To this end, we propose a task-aware parameter-efficient fine-tuning method based on Fisher-whitened Cross-Covariance Analysis (FCCA). By constructing AC-FCCA and AR-FCCA modules, our approach introduces asymmetric coupling and adaptive rank allocation mechanisms to dynamically optimize subspace adaptation capacity under a fixed parameter budget. Experimental results demonstrate that the proposed method significantly outperforms parameter-matched LoRA baselines on both Whisper and Qwen3-ASR, effectively enhancing recognition performance and parameter utilization efficiency in low-resource scenarios.
📝 Abstract
Adapting multilingual speech foundation models to low-resource languages remains difficult, especially for languages that are poorly represented during pre-training. While parameter-efficient fine-tuning (PEFT) reduces the cost of adapting large models, conventional approaches such as LoRA rely on generic low-rank parameterizations and do not explicitly use downstream task information to define the adaptation subspace. To investigate whether task-informed PEFT can better support low-resource ASR, we apply Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR, and introduce two complementary extensions: Asymmetric-Coupled FCCA (AC-FCCA), which exploits structured cross-layer sharing, and Adaptive-Rank FCCA (AR-FCCA), which reallocates adaptation capacity across projection matrices under a fixed parameter budget. Under controlled multilingual experiments, we evaluate these approaches on languages that are poorly represented or unsupported during pre-training alongside well-represented languages. Standard FCCA is competitive with, and usually outperforms, trainable-parameter-budget-matched LoRA. AR-FCCA provides the most consistent improvement over standard FCCA across both model architectures, with statistically significant gains in several evaluation settings, while retaining the same number of trainable parameters. These results show that task-informed subspace construction can be effective for low-resource speech adaptation, and that adaptive rank allocation provides a robust way to improve parameter efficiency without increasing model capacity.
Problem

Research questions and friction points this paper is trying to address.

Low-Resource Speech Recognition
Parameter-Efficient Fine-Tuning
Multilingual Speech Foundation Models
Automatic Speech Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fisher-Whitened Cross-Covariance Analysis
Parameter-Efficient Fine-Tuning
Low-Resource Speech Recognition
Adaptive-Rank Allocation
Cross-Layer Sharing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Asmee Mishra
Department of Engineering, University of Cambridge, United Kingdom
Mengjie Qian
Mengjie Qian
University of Cambridge
speech recognitionmachine learningspoken language assessmentlow-resource
B
Brechtje Post
Phonetics Laboratory, University of Cambridge, United Kingdom
Kate Knill
Kate Knill
University of Cambridge
Speech technologyinteractive technology