PHAST-Net: Attention-Guided, Physics-Informed Network for Unified Estimation of Ideal Time-Frequency Representations

📅 2026-06-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing methods struggle to jointly estimate multiple ideal time–frequency representations—such as spectrograms and beat-synchronous chromagrams—with high resolution and without cross-term interference, particularly for nonstationary signals like speech and music. This work proposes PHAST-Net, an attention-guided, physics-informed neural network that maps a unified input representation, derived from the Continuous Log-Frequency Adaptive Wavelet Transform (CLAWT) scalogram, to diverse target time–frequency representations. Key innovations include constructing the CLAWT scalogram via Cohen’s class kernels, enforcing energy conservation and transform consistency through a physics-informed reprojection loss, and introducing Harmonic PHAST-Net to disentangle fundamental frequency structures alongside Spline-PHAST-Net, which parameterizes time–frequency ridges with spline trajectories to enable reconstruction on arbitrary grids. Experiments demonstrate that the proposed approach significantly outperforms existing techniques on speech, music, and other nonstationary signals, achieving unified, high-fidelity, and cross-term-free estimation of ideal time–frequency representations.
📝 Abstract
We introduce PHAST-Net, an attention-guided, physics-informed network for unified estimation of Ideal Time-Frequency Representations (ITFRs), spanning spectral, tempo-based, metrical, and harmonic representations such as Spectrograms, Tempograms, and Metrograms. PHAST-Net learns an application-general mapping from a constellation of wavelet transforms, the proposed Continuous Log-frequency Adaptive Wavelet Transform (CLAWT), to high-resolution, cross-term-suppressed time-frequency (T-F) representations. The proposed constellation of CLAWTs is selected through Cohen's class kernel analysis to maximise curvature coverage in a logarithmic-frequency T-F plane tailored to harmonic signal structure. PHAST-Net further incorporates a proposed physics-informed auxiliary reprojection loss designed to reconstruct the idealised observed CLAWT constellation from the predicted ITFR and the corresponding Cohen's class kernels during training. This auxiliary objective promotes transform consistency and energy conservation, mitigates pathological target sparsity, and enhances optimisation stability. Attention layers further promote effective cross-term suppression across the input constellation. The log-frequency formulation also enables Harmonic PHAST-Net, which estimates a Harmonic ITFR that isolates fundamental structure, supporting robust fundamental-only representations for speech and music, such as derived fundamental Tempograms and Metrograms. We further introduce Spline-PHAST-Net, which parameterises detected and associated T-F ridges as continuous spline trajectories, enabling arbitrary-grid re-rendering and signal reconstruction. Trained on an effectively unbounded procedurally generated dataset, PHAST-Net demonstrates improved accuracy over established approaches, providing a unified framework for high-resolution, cross-term-robust analysis of speech, music, and broader nonstationary signals.
Problem

Research questions and friction points this paper is trying to address.

Ideal Time-Frequency Representations
cross-term suppression
nonstationary signals
harmonic structure
time-frequency analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

PHAST-Net
physics-informed learning
Ideal Time-Frequency Representations
CLAWT
attention mechanism
J
James M. Cozens
Engineering Department, University of Cambridge, CB2 1PZ Cambridge, U.K.
S
Simon J. Godsill
Engineering Department, University of Cambridge, CB2 1PZ Cambridge, U.K.