Unsupervised Instantaneous Phase and Frequency Tracking by Inverse Voice Synthesis

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the difficulty neural vocoders face in reliably learning fundamental frequency (F0) end-to-end due to weak spectral supervision and missing phase information. To overcome this, we propose an unsupervised approach based on a source-filter architecture that explicitly models instantaneous phase using an anti-aliased additive source and extracts F0 directly through differentiation. The harmonic pathway is learned solely via a combined waveform and spectral loss, eliminating the need for external F0 labels or pitch trackers. This work achieves fully unsupervised, precise recovery of instantaneous phase and frequency. The reconstructed signals attain a signal-to-noise ratio of 8.1 dB, while F0 accuracy surpasses existing fully supervised methods. Furthermore, glottal closure instant detection rates approach REAPER standards, demonstrating highly robust speech synthesis and frequency tracking capabilities.
πŸ“ Abstract
Knowledge-driven neural vocoders struggle to learn reliable fundamental frequency end-to-end, because spectral objectives provide weak supervision of periodic structure and lack phase information. We address this with a source-filter model whose alias-free additive source makes the instantaneous phase of the glottal cycle explicit; differentiating it yields the instantaneous frequency, and thus $F_0$, without an external tracker. Waveform error supervises only the deterministic harmonic path, while a spectral loss covers the full signal. On M4Singer and LM-SSD, the reconstruction is phase-aligned, reaching a signal-to-reconstruction-error ratio of 8.1 dB, while neural baselines remain negative. However, GOLF, given an external $F_0$, still reaches lower spectral distortion. On LM-SSD, the recovered $F_0$ attains the highest overall accuracy of any method tested, including supervised neural pitch trackers applied off the shelf, and the glottal closure instants come within 0.53 points of REAPER's identification rate, without any $F_0$ label.
Problem

Research questions and friction points this paper is trying to address.

fundamental frequency estimation
neural vocoder
instantaneous phase tracking
unsupervised learning
source-filter model
Innovation

Methods, ideas, or system contributions that make the work stand out.

unsupervised pitch tracking
source-filter model
instantaneous phase
neural vocoder
inverse voice synthesis
πŸ”Ž Similar Papers
No similar papers found.