time-frequency modeling

Designs, implements, and evaluates representations, transforms, and parametric models that jointly describe a signal's temporal evolution and spectral content (spectro-temporal or time-frequency representations such as spectrograms, wavelet/S-transform variants, and time-varying spectral models). These artifacts are used to capture transient and oscillatory structure, improve reconstruction or generation of high-frequency components, and enable analysis tasks such as denoising, feature extraction, segmentation, and synthesis.

time-frequencymodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study systematically investigates the frequency-domain encoding capabilities of the Chronos foundation model, addressing a critical gap in understanding how such models represent fundamental signal properties. Through controlled experiments using discrete sinusoidal signals and a lightweight online Minimum Description Length (MDL) probing framework, the work examines the existence, separability, and cross-spectral fidelity of internal frequency representations within the Chronos decoder. The research reveals, for the first time, a degradation in representation quality in high-frequency regions, thereby delineating both the strengths and limitations of Chronos’s frequency encoding mechanism. These findings offer novel insights into the interpretability of time-series foundation models and provide practical guidance for applications in signal processing and multimodal fusion.

foundation modelsfrequency representationmodel interpretability

Beyond the Time Domain: Recent Advances on Frequency Transforms in Time Series Analysis

Feb 12, 2025
QZ
Qianru Zhang
🏛️ The University of Hong Kong | University of California | University of Maryland | Aalborg University | Microsoft | Westlake University | The University of Queensland

Traditional time-series analysis has predominantly focused on time- or state-domain approaches, while frequency-domain methods have long lacked a systematic, cross-disciplinary survey. Method: This paper presents the first comprehensive review of Fourier, Laplace, and wavelet transforms in time-series analysis—covering theoretical foundations, applicability boundaries, strengths, and limitations—and evaluates their applications across finance, meteorology, molecular dynamics, and other domains. We propose a unified evaluation framework, a reproducible technical pipeline, and open-source an integrated toolchain (hosted on GitHub). Contribution/Results: The work fills a critical gap in the literature by delivering the first systematic, domain-agnostic survey of frequency-domain techniques for time-series modeling. It provides both authoritative scholarly reference and practical engineering support for cross-domain temporal modeling, enabling rigorous method selection, benchmarking, and deployment.

Compare strengths and limitations of Fourier, Laplace, Wavelet TransformsExplore applications in finance, molecular, weather, and other fieldsReview frequency transform techniques in time series analysis

This study addresses the challenge of extracting secondary modulation information—such as respiratory signals—from nonsinusoidal multicomponent biomedical signals like photoplethysmography, where coupling between oscillatory amplitude envelopes and instantaneous frequencies complicates analysis. To tackle this, the authors propose the TETRIS framework, which is grounded in a generalized adaptive non-harmonic model. TETRIS introduces, for the first time, a data-driven time–frequency plane tiling strategy guided by the instantaneous frequency of the cardiac component, combined with second-order oscillatory dynamic modeling to adaptively process distinct regions of the signal. This joint approach enables simultaneous decoding of nested rhythms embedded in both amplitude modulation and instantaneous frequency. Experimental results demonstrate that TETRIS significantly enhances time–frequency representation quality on semi-synthetic signals and accurately reconstructs multiple surrogate respiratory signals with high precision.

amplitude modulationinstantaneous frequencyoscillatory dynamics

Point Processes and spatial statistics in time-frequency analysis

Feb 29, 2024
BP
Barbara Pascal
🏛️ Nantes Université | Univ. Lille

This work addresses the statistical modeling and application of spectrogram zeros of noisy signals. Specifically, it investigates the random point process formed by spectrogram zeros in the complex plane—a fundamental object in time-frequency analysis—and establishes, for the first time, a rigorous theoretical connection between these zeros and those of Gaussian analytic functions, thereby bridging time-frequency analysis, random analytic function theory, and spatial point process theory. Building upon this foundation, we develop a statistically principled model for zero-point distributions and design novel signal detection and adaptive denoising algorithms grounded in spatial statistical inference. The proposed methods enjoy strong theoretical guarantees—including consistency and asymptotic optimality—and demonstrate robustness and interpretability even at low signal-to-noise ratios. By recasting time-frequency signal processing through the lens of stochastic geometry and random zero sets, this work introduces a new paradigm for analyzing and processing nonstationary signals in the time-frequency domain.

Analyzing time-frequency content of signals using spectrogramsDeveloping signal detection and denoising algorithmsStudying zeros of spectrograms as Point Processes

Unsupervised Multi-modal Feature Alignment for Time Series Representation Learning

Dec 09, 2023
CL
Chen Liang
🏛️ Harbin Institute of Technology

To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.

Aligning multi-modal time series features for representation learningImproving inductive bias in unsupervised time series encodingReducing feature fusion complexity to enhance scalability

Latest Papers

What's happening recently
View more

This study investigates how to select spectrogram representations that align with the architecture of downstream classifiers to enhance performance in audio and speech analysis tasks. By systematically exploring the design space of spectrograms—encompassing time–frequency resolution, temporal span, and element-wise scaling—and integrating convolutional neural networks with time–frequency analysis techniques, the work comprehensively evaluates diverse spectrogram configurations across multiple tasks. The findings reveal key principles for the co-optimization of front-end features and back-end models, delineate the suitability of various spectrogram representations for specific scenarios, and provide both theoretical grounding and practical guidance for task-oriented, efficient feature engineering and model design.

audio analysisclassifier architecturefeature representation

This work addresses the challenges of feature shift, temporal drift, and spectral discrepancy in source-free domain adaptation for time series. To tackle these issues without access to source-domain data, the authors propose a time–frequency joint alignment approach that explicitly corrects spectral shifts—a first in the source-free setting. By modeling the source domain’s temporal dependencies and spectral characteristics at multiple scales, they design a trainable frequency-domain adaptation module that modulates both phase and amplitude of target-domain signals to achieve distribution alignment. Extensive experiments demonstrate that the proposed method significantly outperforms existing source-free time series adaptation techniques across multiple benchmark datasets, confirming its effectiveness and robustness.

feature shiftsource-free domain adaptationspectral shift

This work addresses the fundamental trade-off between time and frequency resolution inherent in conventional time–frequency representations such as the short-time Fourier transform (STFT), which is governed by the Gabor–Heisenberg uncertainty principle. To overcome this limitation, the authors propose a super-resolution spectrogram fusion method based on optimal transport (OT). By designing a novel transport cost function that preserves time–frequency geometric structure while remaining computationally efficient, they formulate a block majorization–minimization algorithm within an unbalanced OT framework. This approach enables flexible fusion of spectrograms defined on arbitrary time–frequency grids without requiring a common grid alignment, effectively integrating their locally optimal time–frequency characteristics. Experiments on both synthetic signals and real-world speech demonstrate that the proposed method consistently outperforms existing unsupervised fusion techniques in both qualitative and quantitative evaluations.

non-stationary signalsspectrogram fusionsuper-resolution

Existing methods struggle to jointly estimate multiple ideal time–frequency representations—such as spectrograms and beat-synchronous chromagrams—with high resolution and without cross-term interference, particularly for nonstationary signals like speech and music. This work proposes PHAST-Net, an attention-guided, physics-informed neural network that maps a unified input representation, derived from the Continuous Log-Frequency Adaptive Wavelet Transform (CLAWT) scalogram, to diverse target time–frequency representations. Key innovations include constructing the CLAWT scalogram via Cohen’s class kernels, enforcing energy conservation and transform consistency through a physics-informed reprojection loss, and introducing Harmonic PHAST-Net to disentangle fundamental frequency structures alongside Spline-PHAST-Net, which parameterizes time–frequency ridges with spline trajectories to enable reconstruction on arbitrary grids. Experiments demonstrate that the proposed approach significantly outperforms existing techniques on speech, music, and other nonstationary signals, achieving unified, high-fidelity, and cross-term-free estimation of ideal time–frequency representations.

cross-term suppressionharmonic structureIdeal Time-Frequency Representations

This work proposes BandTok, a generative two-dimensional tokenizer for Mel-spectrogram representation that addresses the limitations of existing high-fidelity music codecs relying on residual vector quantization. Such approaches suffer from strong sequential dependencies after flattening, which hinder autoregressive modeling and exacerbate error accumulation. In contrast, BandTok assigns discrete tokens to individual Mel frequency bands per time frame using a single shared codebook, yielding a physically interpretable and structurally disentangled time–frequency token grid. Coupled with a 2D Rotary Position Embedding–enhanced autoregressive language model, this method formulates music generation as image-like time–frequency modeling. Evaluated under data-constrained conditions, BandTok significantly outperforms residual codebook–based baselines in both reconstruction fidelity and controllable generation. Code and audio samples are publicly released.

audio tokenizerautoregressive music generationerror accumulation

Hot Scholars

HT

Hau-Tieng Wu

Courant Institute of Mathematical Sciences, New York University
Mathematical foundation of data analysis. Medical signal processing. High dimensional statistics.
KY

Kohei Yatabe

Tokyo University of Agriculture and Technology
SK

Saif Khan Mohammed

Professor of Electrical Engineering, I. I. T. Delhi
Zak-OTFS waveforms for 6GMassive MIMO systemsLarge MIMO systemsLarge scale antenna systems
YT

Yu Tsao

Research Fellow (Professor), Deputy Director, CITI, Academia Sinica
Assistive Oral Communication TechnologiesSpeech EnhancementVoice ConversionSpeech Assessment
MF

Mahmoud Fakhry

Universidad Francisco de Vitoria
Signal AnalysisNeural NetworksArtificial Intelligence