Score
Designs, implements, and evaluates automatic speech recognition systems that use state-space models (SSMs) such as Mamba as their sequence model, including SSM-based encoder/decoder architectures, training pipelines, and inference components. Builds and analyzes methods to improve efficiency and scalability (reduced compute and training time) and to handle long-range temporal dependencies and generalization to longer input sequences.
Structured State Space Models (SSMs) face a fundamental trade-off between long-range dependency modeling and computational efficiency, limiting their broad adoption across NLP, speech, vision, and time-series domains. This paper provides the first systematic survey of SSMs—from theoretical foundations (continuous-time dynamics, HiPPO projections) to industrial variants (S4, Mamba, S5, Jamba)—unifying analysis of their linear-time complexity, memory-efficient parameterization, and hardware-aware inference acceleration. We identify selectivity mechanisms and low-rank structured matrices as key innovations enabling SSMs to emerge as the third major sequence modeling paradigm—alongside RNNs and Transformers—achieving near-Transformer accuracy on long-sequence tasks while reducing memory footprint by over 70% and significantly improving inference throughput. We further highlight critical open challenges: training instability, hybrid modeling strategies, and interpretability.
This work addresses speech enhancement (SE) by introducing Mamba—a non-attention, scalable state-space model (SSM)—to end-to-end regression modeling for the first time, proposing the SEMamba architecture supporting both causal and non-causal configurations. To improve perceptual quality, we design a perceptual contrastive stretching (PCS) module, jointly optimized with signal-level and metric-oriented losses. Evaluated on the VoiceBank-DEMAND benchmark, SEMamba achieves a new state-of-the-art PESQ score of 3.69, reducing FLOPs by approximately 12% compared to leading Transformer-based methods. Moreover, as an ASR front-end, it demonstrates competitive performance. This study validates the effectiveness of SSMs for speech temporal modeling and establishes a novel paradigm for lightweight, efficient, and high-fidelity speech enhancement.
This study systematically clarifies the similarities, differences, and theoretical connections between state space models (SSMs) and hidden Markov models (HMMs) in sequence modeling. By employing a unified probabilistic graphical model framework, it compares classical probabilistic SSMs—including linear Gaussian SSMs and Kalman filtering—with HMMs and modern neural SSMs, analyzing their structural properties, inference algorithms, and learning mechanisms. The work identifies precise conditions under which these models are equivalent or fundamentally distinct, establishing formal correspondences across representation, inference, and training paradigms. In doing so, it provides a principled probabilistic interpretation of modern neural SSMs, bridges conceptual gaps among control theory, probabilistic modeling, and deep learning, and offers theoretical guidance for informed model selection and design.
This work investigates the fundamental expressive capacity of linear state-space models (SSMs) for language modeling, clarifying their theoretical modeling boundaries relative to Transformers and classical RNNs. Method: Leveraging formal language theory and automata theory, we formally characterize SSM expressivity—proving for the first time that linear SSMs can exactly recognize star-free languages and optimally model bounded hierarchical structures in memory. We identify a critical expressivity bottleneck in contemporary SSM designs arising from the absence of nonlinearity in state updates. Contribution/Results: Our analysis reveals that SSMs and Transformers possess complementary—not substitutive—capabilities. Empirical evaluation on the Mamba architecture demonstrates substantial gains over Transformers on star-free language tasks and superior memory efficiency in hierarchical structure modeling. These findings provide both theoretical foundations and practical guidance for designing next-generation efficient large language model architectures.
Long temporal dependencies and strong semantic correlations in speech signals pose challenges for efficient contextual modeling, motivating exploration of alternatives to Transformer self-attention. Method: This work investigates the applicability of selective state space models (SSMs), specifically Mamba, to speech processing, proposing Bidirectional Mamba (BiMamba)—a novel architecture that enhances context capture via bidirectional state propagation—and integrating it into end-to-end automatic speech recognition (ASR) and speech enhancement frameworks, accompanied by ablation-driven architectural adaptation strategies. Contribution/Results: This is the first systematic empirical validation of Mamba-style SSMs in speech tasks. BiMamba consistently outperforms both standard Mamba and Transformer baselines on ASR and speech enhancement benchmarks, with particularly pronounced gains in semantically sensitive scenarios. Results demonstrate BiMamba’s superior efficiency, generalization capability, and viability as a scalable, attention-free paradigm for speech modeling.
This work investigates whether structured state space models require complex input-dependent mechanisms for multivariate time series classification. Through systematic evaluation of diagonal state space models (S4D) against Mamba-style input-dependent architectures, the study demonstrates that simplified designs can achieve competitive performance. To this end, the authors propose MS4 and its normalized variant MS4N—lightweight models incorporating only linear input projection and channel mixing. Extensive experiments across 59 datasets from MONSTER and UEA benchmarks show that MS4N outperforms Mamba in both accuracy and efficiency while using fewer parameters, and matches or exceeds the performance of state-of-the-art deep learning models that are 2–10 times larger. These results underscore the efficacy and superiority of minimalist architectural design in this domain.
This work investigates the selective forgetting mechanism of state space models (SSMs)—particularly the Mamba family (130M–1.4B parameters)—under fixed memory constraints when processing long sequences. To address *which semantic types and sequences are more prone to forgetting*, we propose an autoencoder-based latent-state reconstruction evaluation framework that quantifies information loss across token categories (e.g., part-of-speech tags, named entities) and sequence domains (e.g., code, mathematical problems). Our systematic analysis reveals, for the first time, that low-frequency tokens—including mathematical symbols, organizational named entities, and non-standard American English—exhibit significantly higher forgetting rates; crucially, forgetting magnitude is strongly negatively correlated with token frequency in the pretraining corpus. This establishes an interpretable, data-distribution-aware linkage between forgetting patterns and training statistics, providing both empirical grounding and a diagnostic tool for memory modeling and long-context optimization in SSMs.
This work addresses the trade-off between inference efficiency and modeling capacity in existing linear sequence models, which often sacrifice state tracking performance and struggle to realize theoretical linear complexity on real hardware. From an inference-first perspective, the paper introduces three key innovations within the state space model (SSM) framework: a highly expressive recurrence derived from SSM discretization, a complex-valued state update mechanism enabling richer dynamics, and a multi-input multi-output (MIMO) architecture that incurs no decoding latency. Evaluated at 1.5B parameters, the proposed model outperforms the strongest baseline, Gated DeltaNet, by 0.6 accuracy points on average, with the MIMO variant yielding an additional 1.2-point gain (total +1.8%). Notably, it matches Mamba-2’s perplexity using only half the state size, substantially advancing the Pareto frontier of performance versus efficiency.
Systematic understanding of runtime behavior, resource consumption, and scalability of selective state space models (SSMs) remains lacking. Method: This paper conducts multi-granularity performance profiling of Mamba-1/2, quantitatively characterizing SSM computational patterns, memory access bottlenecks, and state activity distributions across varying sequence lengths. Based on these insights, we propose a state-activity-driven structured pruning method that dynamically prunes low-activity states during inference to jointly optimize accuracy and efficiency. Contribution/Results: Experiments demonstrate an average 1.14× throughput speedup and 11.50% memory compression across diverse sequence lengths—substantially outperforming baselines. This work establishes a reproducible empirical foundation and a novel paradigm for lightweight SSM design and hardware-aware optimization.
Multilingual automatic speech recognition (ASR) suffers from performance disparity between high-resource and low-resource languages. To address this, we propose the first integration of the Mamba state space model into multilingual ASR, replacing the conventional Transformer architecture. Leveraging Mamba’s linear-time complexity and superior long-range dependency modeling, our approach establishes an efficient, scalable, unified ASR framework. Crucially, it incorporates an implicit language-aware mechanism and shared cross-lingual representations to significantly improve modeling of low-resource languages. Evaluated on standard multilingual benchmarks—including MLS and CommonVoice—our method achieves competitive word error rates relative to state-of-the-art Transformer-based models, while accelerating inference by approximately 2.3× and reducing GPU memory consumption by 40%. This work introduces a novel paradigm for efficient and equitable multilingual ASR, advancing both computational efficiency and linguistic fairness.