🤖 AI Summary
To address the challenges of long-range contextual dependencies and language-specific constraints in universal phoneme recognition, this paper proposes a context-agnostic and language-agnostic universal phoneme encoding method. The approach extracts acoustic features from fixed-width, 120-ms short windows independently—eliminating reliance on sequential modeling. A lightweight single-window model architecture is designed to explicitly capture cross-lingual commonalities in acoustic patterns. Supervised and self-supervised learning objectives are jointly optimized across multilingual data, enabling zero-shot cross-lingual transfer. Evaluated on multilingual phoneme recognition benchmarks, the method achieves competitive performance. Notably, it significantly outperforms existing context-free approaches on the UCLA zero-shot cross-lingual evaluation, demonstrating superior generalization. This work is the first to empirically validate the feasibility and strong transferability of short-window independent encoding for universal phoneme representation learning.
📝 Abstract
Universal phoneme recognition typically requires analyzing long speech segments and language-specific patterns. Many speech processing tasks require pure phoneme representations free from contextual influence, which motivated our development of CUPE - a lightweight model that captures key phoneme features in just 120 milliseconds, about one phoneme's length. CUPE processes short, fixed-width windows independently and, despite fewer parameters than current approaches, achieves competitive cross-lingual performance by learning fundamental acoustic patterns common to all languages. Our extensive evaluation through supervised and self-supervised training on diverse languages, including zero-shot tests on the UCLA Phonetic Corpus, demonstrates strong cross-lingual generalization and reveals that effective universal speech processing is possible through modeling basic acoustic patterns within phoneme-length windows.