Quality-Aware Self-Correcting Speech Translation on an Edge Device

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of correcting low-quality translations under resource constraints in offline speech translation for edge devices. We propose a fully offline self-correction pipeline deployed on a Jetson Nano, integrating lightweight modules including Whisper-tiny, Opus-MT, and Minimum Bayes Risk decoding. Methodologically, a quality estimation (QE) gate triggers secondary refinement to enhance translation quality without retraining. Notably, we observe that QE fails as an effective reranker and consequently exclude it from candidate selection, thereby conserving memory along the critical path. Experimental results on FLORES-200 demonstrate significant improvements in BLEU, ChrF, and COMET metrics. Furthermore, the proposed approach frees 680MB of memory and enables real-time demonstrations across six language pairs.
📝 Abstract
We present a fully offline speech-to-speech translation pipeline that runs on a Jetson Nano (4 GB) and corrects its own weak translations without retraining. A Whisper-tiny ASR feeds an Opus-MT translator; multilingual BERT cosine similarity acts as a Quality Estimation (QE) gate, triggering a secondary-pass correction when confidence falls below a pre-defined threshold $τ$. We compare three correction methods: QE reranking (M1), Minimum Bayes-Risk decoding (M2), and constrained beam search (M3). On 1,012 FLORES-200 sentences (English-Spanish), M2 at $τ=0.90$ produces statistically significant improvements over greedy decoding on BLEU (+0.67, p<0.001), ChrF (+0.51, p<0.001), and COMET (+0.0020 at N=3, p=0.002); M1 yields no significant gains, and M3 is significantly worse than baseline (p>0.99). Our central finding is that QE functions effectively as a gate but poorly as a ranker: removing the QE model from candidate selection (M1$\to$M2) does not hurt quality and frees 680 MB from the critical path. Using a gain-to-edit ratio adapted from the post-editing-effort literature, we further show that smaller candidate pools (N=3) yield more surgical corrections with better semantic adequacy, while larger pools (N=10) maximise lexical reward. We release the system and demonstrate live translation across six language pairs.
Problem

Research questions and friction points this paper is trying to address.

speech translation
edge device
quality estimation
self-correction
offline pipeline
Innovation

Methods, ideas, or system contributions that make the work stand out.

Edge Device Translation
Quality Estimation
Self-Correcting
Minimum Bayes-Risk Decoding
Offline Pipeline
💼 Related Jobs
No related jobs found.