Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of multi-pitch estimation in vocal ensembles, where overlapping fundamental frequencies complicate analysis and existing methods incur high feature extraction costs. We systematically compare the representational efficacy of the linear short-time Fourier transform (STFT) and the harmonic constant-Q transform (HCQT). Challenging the conventional assumption that high frequency resolution is indispensable, this work validates the effectiveness of fixed frequency resolution for handling time-varying pitches. Experimental results demonstrate that a short-window linear STFT outperforms the HCQT while significantly reducing computational overhead. These findings confirm that finer frequency resolution is not a critical factor for improving multi-pitch estimation performance, offering a more efficient alternative for polyphonic audio analysis.
📝 Abstract
Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.
Problem

Research questions and friction points this paper is trying to address.

multi-pitch estimation
vocal ensembles
time-frequency representations
harmonic constant-Q transform
short-time Fourier transform
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-pitch estimation
Vocal ensembles
Linear STFT
Harmonic CQT
Time-frequency representations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.