SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current wearable silent speech interfaces are limited by closed vocabularies and reliance on invasive hardware. This work proposes the first large-scale, open-vocabulary silent speech dataset based on non-invasive acoustic-sensing eyewear, simultaneously capturing three modalities: ultrasonic echoes, vocal audio (for alignment), and frontal video. The dataset comprises 34 hours of recordings with 18,000 utterances, covering the full phonetic inventory and contemporary conversational English. Using a ResNet-34 architecture trained with CTC loss as the baseline model, the study achieves a word error rate (WER) of 26.3% on open-vocabulary silent speech recognition, establishing the first public benchmark for this task.
📝 Abstract
Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-scale, open-vocabulary, trimodal dataset for wearable SSI using acoustic-sensing eyewear. It contains 34 hours across 18,000 utterances with three synchronized modalities: ultrasound echo profiles, voiced audio, and frontal video, in both voiced and silent modes. The corpus draws from the SODA dialogue dataset, providing contemporary conversational English with 5,356 unique words and full phoneme coverage. A CTC-based ResNet-34 baseline achieves 26.3% word error rate (WER) on open-vocabulary silent speech recognition, the first benchmark for this task. Dataset is available at https://doi.org/10.7298/xjjr-9m85
Problem

Research questions and friction points this paper is trying to address.

silent speech interface
open-vocabulary
wearable
large-scale dataset
tri-modal
Innovation

Methods, ideas, or system contributions that make the work stand out.

silent speech interface
open-vocabulary
trimodal dataset
acoustic-sensing eyewear
speech recognition
🔎 Similar Papers
No similar papers found.