instantaneous pitch estimation

Designs and implements algorithms and analysis pipelines that estimate instantaneous fundamental frequency (F0) or pitch as a time-varying curve from audio or other pitched time-series by computing analytic signals and tracking phase evolution to derive instantaneous frequency. These methods produce high-resolution, time-varying pitch tracks (pitch tracking/instantaneous frequency estimation) and are engineered to remain accurate under rapid pitch changes and degraded or noisy inputs.

instantaneouspitchestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the difficulty neural vocoders face in reliably learning fundamental frequency (F0) end-to-end due to weak spectral supervision and missing phase information. To overcome this, we propose an unsupervised approach based on a source-filter architecture that explicitly models instantaneous phase using an anti-aliased additive source and extracts F0 directly through differentiation. The harmonic pathway is learned solely via a combined waveform and spectral loss, eliminating the need for external F0 labels or pitch trackers. This work achieves fully unsupervised, precise recovery of instantaneous phase and frequency. The reconstructed signals attain a signal-to-noise ratio of 8.1 dB, while F0 accuracy surpasses existing fully supervised methods. Furthermore, glottal closure instant detection rates approach REAPER standards, demonstrating highly robust speech synthesis and frequency tracking capabilities.

fundamental frequency estimationinstantaneous phase trackingneural vocoder

This study addresses the limitations of traditional fundamental frequency extraction methods, which suffer from filtering-induced distortions in harmonic- and noise-contaminated speech, leading to insufficient accuracy in instantaneous pitch estimation. The work proposes a novel approach by formulating fundamental frequency extraction as a speech enhancement task and introduces an end-to-end framework based on Wave-U-Net to directly estimate the fundamental waveform from noisy speech. Instantaneous pitch is then derived by computing the instantaneous frequency of the estimated waveform using analytic signal theory. Evaluated across diverse scenarios—including speech, singing, musical instruments, and degraded speech—the proposed method significantly outperforms conventional deterministic approaches, achieving higher estimation accuracy and enhanced robustness.

fundamental waveformharmonicsinstantaneous pitch estimation

A Robust Method for Pitch Tracking in the Frequency Following Response using Harmonic Amplitude Summation Filterbank

Jun 23, 2025
SS
Sajad Sadeghkhani
🏛️ University of Ottawa | Shahid Bahonar University of Kerman

This paper addresses inaccurate fundamental frequency (F₀) tracking in frequency-following responses (FFRs). We propose a robust harmonic-structure-based F₀ estimation method. Our approach introduces a stimulus-aware harmonic amplitude summation (HAS) filterbank that enhances the F₀ and its integer harmonics while suppressing non-harmonic noise. Instead of conventional autocorrelation-based peak detection, we employ frequency-domain amplitude aggregation combined with a most-prominent-peak criterion, constrained by the known stimulus F₀. This work is the first to systematically integrate harmonic prior knowledge into the FFR F₀ estimation framework. Evaluated on FFR data from 16 subjects elicited by four natural speech stimuli, our method reduces root-mean-square error (RMSE) in F₀ tracking by 8.8%–47.4% relative to the classical autocorrelation method, significantly improving dynamic pitch representation accuracy.

Enhancing accuracy of F0 extraction compared to Autocorrelation Function method.Improving F0 tracking in Frequency Following Response using harmonic structure.Reducing noise in FFR pitch estimation with stimulus-aware filterbank.

Improving Neural Pitch Estimation with SWIPE Kernels

Jul 15, 2025
DM
David Marttila
🏛️ Queen Mary University of London

This study addresses the longstanding challenge in neural pitch estimation of simultaneously achieving high accuracy, robustness, and parameter efficiency. We propose embedding task-specific, handcrafted features—specifically, the sawtooth-wave-inspired SWIPE kernel—into the front-end of neural networks. Crucially, we reformulate the classical SWIPE algorithm as a differentiable, learnable audio preprocessing module, replacing conventional spectral representations for the first time. Experiments demonstrate three key contributions: (1) Integrating the SWIPE front-end reduces model parameter count by an order of magnitude while maintaining or improving performance in both supervised and self-supervised settings; (2) SWIPE alone—without any neural backbone—outperforms state-of-the-art self-supervised neural pitch estimators, revealing its inherent discriminative power; and (3) the approach significantly enhances noise robustness. This work establishes a new paradigm for lightweight, highly robust pitch modeling by unifying principled signal processing with deep learning.

Enhancing noise robustness in pitch estimation using task-specific featuresImproving neural pitch estimation accuracy with SWIPE kernelsReducing neural network size without performance loss via SWIPE frontend

Point Processes and spatial statistics in time-frequency analysis

Feb 29, 2024
BP
Barbara Pascal
🏛️ Nantes Université | Univ. Lille

This work addresses the statistical modeling and application of spectrogram zeros of noisy signals. Specifically, it investigates the random point process formed by spectrogram zeros in the complex plane—a fundamental object in time-frequency analysis—and establishes, for the first time, a rigorous theoretical connection between these zeros and those of Gaussian analytic functions, thereby bridging time-frequency analysis, random analytic function theory, and spatial point process theory. Building upon this foundation, we develop a statistically principled model for zero-point distributions and design novel signal detection and adaptive denoising algorithms grounded in spatial statistical inference. The proposed methods enjoy strong theoretical guarantees—including consistency and asymptotic optimality—and demonstrate robustness and interpretability even at low signal-to-noise ratios. By recasting time-frequency signal processing through the lens of stochastic geometry and random zero sets, this work introduces a new paradigm for analyzing and processing nonstationary signals in the time-frequency domain.

Analyzing time-frequency content of signals using spectrogramsDeveloping signal detection and denoising algorithmsStudying zeros of spectrograms as Point Processes

Latest Papers

What's happening recently
View more

This study addresses the challenge of multi-pitch estimation in vocal ensembles, where overlapping fundamental frequencies complicate analysis and existing methods incur high feature extraction costs. We systematically compare the representational efficacy of the linear short-time Fourier transform (STFT) and the harmonic constant-Q transform (HCQT). Challenging the conventional assumption that high frequency resolution is indispensable, this work validates the effectiveness of fixed frequency resolution for handling time-varying pitches. Experimental results demonstrate that a short-window linear STFT outperforms the HCQT while significantly reducing computational overhead. These findings confirm that finer frequency resolution is not a critical factor for improving multi-pitch estimation performance, offering a more efficient alternative for polyphonic audio analysis.

harmonic constant-Q transformmulti-pitch estimationshort-time Fourier transform

This study systematically investigates the frequency-domain encoding capabilities of the Chronos foundation model, addressing a critical gap in understanding how such models represent fundamental signal properties. Through controlled experiments using discrete sinusoidal signals and a lightweight online Minimum Description Length (MDL) probing framework, the work examines the existence, separability, and cross-spectral fidelity of internal frequency representations within the Chronos decoder. The research reveals, for the first time, a degradation in representation quality in high-frequency regions, thereby delineating both the strengths and limitations of Chronos’s frequency encoding mechanism. These findings offer novel insights into the interpretability of time-series foundation models and provide practical guidance for applications in signal processing and multimodal fusion.

foundation modelsfrequency representationmodel interpretability

The scarcity of large-scale, standardized engine audio datasets with precise operating condition annotations has hindered advances in active sound design and data-driven synthesis. This work proposes an analysis-driven procedural generation framework that extracts harmonic structures from real-world recordings via pitch-adaptive spectral analysis and drives an extended parametric harmonic-plus-noise synthesizer, enabling sample-level control over RPM and torque. The resulting Procedural Engine Sounds Dataset comprises 19 hours of audio across 5,935 samples, spanning a wide range of operating conditions and acoustic complexities. This dataset effectively addresses the gap in real-world data availability while preserving authentic harmonic characteristics, thereby supporting learning-based parameter estimation and audio synthesis tasks.

control annotationsdata scarcityengine sound dataset

Traditional Fourier analysis is constrained by periodic boundary conditions, making it ill-suited for high-resolution time–frequency analysis of non-stationary, aperiodic pulmonary sound pulse trains—such as crackles and wheezes. This work proposes replacing the conventional periodicity assumption with linear extrapolation boundary conditions to construct an instantaneous spectral analysis method that circumvents the windowing limitations inherent in short-time Fourier transform. For the first time, this approach enables independent spectral extraction and reconstruction of individual pulses within stochastic pulse sequences. The method substantially enhances time–frequency resolution, successfully visualizing the fine time–frequency structures of both normal and pathological lung sounds and clearly revealing the spectral characteristics of individual acoustic pulses.

Fourier analysisinstantaneous spectralung sounds

This study investigates how music foundation models encode pitch information in their internal representations, with a particular focus on whether octave periodicity is inherently captured. By feeding isolated musical notes into pretrained models and employing principal component analysis, visualization of intermediate representations, and cross-model comparisons, the work provides the first empirical evidence that pitch is embedded within these models as a spiral geometric structure. The research further demonstrates that the clarity and morphology of this spiral are influenced by both model architecture and the acoustic characteristics of the input audio. These findings offer novel insights into the representational mechanisms underlying music foundation models and advance our understanding of how fundamental musical attributes are structured in deep neural networks.

helical structureintermediate representationsmusic foundation models

Hot Scholars

GR

Gaël Richard

Professor, Télécom Paris, Institut polytechnique de Paris
Audio signal processingMachine listeningMusic ProcessingMusic Information Retrieval
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing
HK

Hemant Kumar Kathania

Assistant Professor NIT Sikkim
Children Speech RecognitionLow ResourceZero shotkeyword spotting
PS

Paban Sapkota

Research Scholar (NIT Sikkim)
Speech RecognitionDisordered Speech AnalysisData AugmentationSpeech Feature Engineering
SR

Sudarsana Reddy Kadiri

University of Southern California
Speech ProcessingBiomedical SignalsMultimodalityHealthcare Informatics