ssm-based audio modeling

Designs, implements, and analyzes state‑space sequence models specifically for audio signals, including SSM‑based backbone architectures that act as feature extractors and compact distilled variants derived from larger SSMs. Work focuses on building models and training/distillation procedures that capture long‑range temporal dependencies while preserving mid‑to‑high spectral components and characterizing their spectral/temporal response relative to alternative sequence backbones.

ssm-basedaudiomodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Survey on Structured State Space Sequence (S4) Models

Mar 22, 2025
SS
Shriyank Somvanshi
🏛️ Texas State University

Structured State Space Models (SSMs) face a fundamental trade-off between long-range dependency modeling and computational efficiency, limiting their broad adoption across NLP, speech, vision, and time-series domains. This paper provides the first systematic survey of SSMs—from theoretical foundations (continuous-time dynamics, HiPPO projections) to industrial variants (S4, Mamba, S5, Jamba)—unifying analysis of their linear-time complexity, memory-efficient parameterization, and hardware-aware inference acceleration. We identify selectivity mechanisms and low-rank structured matrices as key innovations enabling SSMs to emerge as the third major sequence modeling paradigm—alongside RNNs and Transformers—achieving near-Transformer accuracy on long-sequence tasks while reducing memory footprint by over 70% and significantly improving inference throughput. We further highlight critical open challenges: training instability, hybrid modeling strategies, and interpretability.

Addressing long-range dependency modeling in sequence tasksImproving computational efficiency over RNNs and TransformersOptimizing memory and inference speed in SSM models

Exploring State-Space-Model based Language Model in Music Generation

Jul 09, 2025
WL
Wei-Jaw Lee
🏛️ National Taiwan University

This work investigates the application of State Space Models (SSMs) to text-to-music generation, introducing the first adaptation of the Mamba architecture as an efficient music decoder. To enable discrete modeling, audio is encoded via Residual Vector Quantization (RVQ), and we empirically find that a single-layer codebook suffices for capturing salient musical semantics. The SiMBA encoder is then refactored into an autoregressive sequence decoder based on the SSM. Compared to Transformer baselines, our SSM decoder achieves faster convergence under low-resource training conditions and significantly more efficient inference, while yielding generated audio with superior fidelity and musicality—closer to real-world recordings. This study demonstrates the viability of SSMs for modeling long-range musical structure under computational constraints, establishing a new paradigm for lightweight, high-fidelity music generation.

Achieve efficient music generation under limited resourcesCompare SiMBA with Transformer decoders in music synthesisExplore Mamba-based models for text-to-music generation

To address the O(L²) computational and memory overhead of Transformer-based audio models caused by self-attention, this paper proposes SSAMBA—the first attention-free Mamba architecture tailored for self-supervised audio representation learning. Methodologically, SSAMBA introduces three key innovations: (1) the first adaptation of the state-space model (Mamba) to audio processing; (2) a bidirectional Mamba structure designed to capture long-range temporal-spectral dynamics; and (3) a unified pretraining framework jointly optimizing discriminative (contrastive learning) and generative (masked spectrogram reconstruction) objectives. Evaluated on audio classification, keyword spotting, speaker identification, and emotion recognition, SSAMBA consistently outperforms SSAST. Its Tiny variant achieves 92.7% inference speedup and reduces GPU memory consumption by 95.4% (at 22K tokens), significantly enhancing efficiency for edge deployment.

Efficient audio representation learningMamba state space model advantagesSelf-supervised attention-free model

Audio xLSTMs: Learning Self-supervised audio representations with xLSTMs

Aug 29, 2024
SY
Sarthak Yadav
🏛️ Aalborg University | Pioneer Centre for Artificial Intelligence | National and Kapodistrian University of Athens

This work investigates the feasibility of extended LSTM (xLSTM) for self-supervised general-purpose audio representation learning. We propose Audio xLSTM (AxLSTM), the first adaptation of xLSTM to this task, leveraging spectrogram patch masking for sequential audio representation learning. AxLSTM preserves strong long-range dependency modeling while substantially improving generalization and parameter efficiency. Pretrained on AudioSet, it outperforms the SSAST baseline by up to 20% in average accuracy across ten downstream audio understanding tasks, with up to 45% fewer parameters. Our key contributions are: (1) pioneering the application of xLSTM to self-supervised audio representation learning; and (2) empirically demonstrating that xLSTM maintains superior sequential modeling capabilities while achieving enhanced transferability to downstream tasks and a more favorable parameter–performance trade-off compared to state-of-the-art alternatives.

Comparing AxLSTM performance with SSAST baselinesEvaluating xLSTM for self-supervised audio representation learningProposing Audio xLSTM for masked spectrogram patch learning

This work investigates the fundamental expressive capacity of linear state-space models (SSMs) for language modeling, clarifying their theoretical modeling boundaries relative to Transformers and classical RNNs. Method: Leveraging formal language theory and automata theory, we formally characterize SSM expressivity—proving for the first time that linear SSMs can exactly recognize star-free languages and optimally model bounded hierarchical structures in memory. We identify a critical expressivity bottleneck in contemporary SSM designs arising from the absence of nonlinearity in state updates. Contribution/Results: Our analysis reveals that SSMs and Transformers possess complementary—not substitutive—capabilities. Empirical evaluation on the Mamba architecture demonstrates substantial gains over Transformers on star-free language tasks and superior memory efficiency in hierarchical structure modeling. These findings provide both theoretical foundations and practical guidance for designing next-generation efficient large language model architectures.

Current SSMs have design limits affecting their expressivenessSSMs handle star-free state tracking better than transformersSSMs' expressive power compared to transformers and RNNs

Latest Papers

What's happening recently
View more

This work addresses the challenge that traditional time-invariant models struggle to effectively capture switching dynamics in time-varying systems. To overcome this limitation, the authors propose a neural network–based time-varying state-space model that incorporates a learnable dictionary of time-varying basis functions. This design flexibly represents diverse temporal evolution patterns of system dynamics while maintaining manageable computational complexity, substantially enhancing the model’s capacity to capture switching sequences. The study further reveals an optimal allocation strategy for time-varying degrees of freedom across model components. Experimental results demonstrate that the proposed model consistently outperforms existing time-invariant approaches on both synthetic switching systems and speech denoising tasks, confirming its effectiveness and strong generalization capability.

signal processingstate-space modelsswitching dynamics

This work investigates whether structured state space models require complex input-dependent mechanisms for multivariate time series classification. Through systematic evaluation of diagonal state space models (S4D) against Mamba-style input-dependent architectures, the study demonstrates that simplified designs can achieve competitive performance. To this end, the authors propose MS4 and its normalized variant MS4N—lightweight models incorporating only linear input projection and channel mixing. Extensive experiments across 59 datasets from MONSTER and UEA benchmarks show that MS4N outperforms Mamba in both accuracy and efficiency while using fewer parameters, and matches or exceeds the performance of state-of-the-art deep learning models that are 2–10 times larger. These results underscore the efficacy and superiority of minimalist architectural design in this domain.

Mambamodel complexitymultivariate time series classification

This work proposes an end-to-end, feature-free audio classification approach based on a parallel deep reservoir computing architecture that operates directly on raw audio waveforms, eliminating the need for explicit feature extraction such as MFCCs. Traditional methods relying on handcrafted features often incur high computational overhead and complex preprocessing pipelines. To evaluate the efficacy of the proposed design, the authors conduct comparative experiments using shallow, serial, and parallel deep reservoir models. Results demonstrate that the parallel architecture achieves significantly superior performance over baseline methods while maintaining low model complexity. The approach enables efficient temporal modeling and hierarchical representation learning, highlighting its scalability and practical potential for audio processing tasks.

acoustic signal preprocessingend-to-end classificationfeature-free

This work explores two underexplored design dimensions of state space models (SSMs) for time series classification: deep recurrence and input reshaping. By recurrently applying the same SSM module across depth (i.e., deep recurrence) and integrating temporal concatenation or feature-time rechunking strategies at the input stage, the proposed approach substantially enhances model performance. The study introduces deep recurrence into SSMs for the first time, revealing its role as an effective inductive bias, and systematically demonstrates the consistent benefits of input reshaping across both low- and high-dimensional time series. Evaluated on six benchmarks, the method matches or surpasses existing large models with fewer parameters, achieving accuracy gains of 1–6%. The combined effect of both techniques is additive and consistently validated across multiple random seeds.

depth-recurrenceinput reshapingparameter sharing

Hot Scholars

CP

Changhao Pan

Zhejiang University
Multi-Modal Genarative AISinging Voice Synthesis
YZ

Yu Zhang

ByteDance
Spatial AudioSinging Voice SynthesisMusic GenerationSpeech Synthesis
XY

Xiang Yin

Bytedance AI Lab
video/audio generationspeech-driven avatar
KL

Ke Lei

Zhejiang University
AI generated content
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing