🤖 AI Summary
This study addresses the challenge of multi-pitch estimation in vocal ensembles, where overlapping fundamental frequencies complicate analysis and existing methods incur high feature extraction costs. We systematically compare the representational efficacy of the linear short-time Fourier transform (STFT) and the harmonic constant-Q transform (HCQT). Challenging the conventional assumption that high frequency resolution is indispensable, this work validates the effectiveness of fixed frequency resolution for handling time-varying pitches. Experimental results demonstrate that a short-window linear STFT outperforms the HCQT while significantly reducing computational overhead. These findings confirm that finer frequency resolution is not a critical factor for improving multi-pitch estimation performance, offering a more efficient alternative for polyphonic audio analysis.
📝 Abstract
Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.