Score
Designs and implements state‑space model layers or modules that aggregate and propagate long‑range dependencies along the depth (slice/volume) axis of data, providing efficient depth‑wise context modeling. Integrates these state‑space depth modules into hybrid 2D/3D neural architectures and analyzes their effects on segmentation stability, variance, and overall performance.
This work systematically investigates the effectiveness and efficiency of State Space Models (SSMs) for long-sequence modeling, benchmarking them against Transformers. To address the lack of a unified theoretical framework, we propose the first comprehensive taxonomy of SSM evolution—categorizing mainstream paradigms into classical SSMs, structured SSMs (e.g., S4), and selective SSMs (e.g., Mamba)—and identify three core mechanisms driving performance gains: HiPPO-theory-backed linear time-invariant dynamics, low-rank structured parameterization, and hardware-aware selective scanning. Integrating motivations, mathematical formulations, paradigm comparisons, and representative applications, we construct the first holistic, hierarchically organized SSM knowledge graph. This synthesis fills a critical gap in systematic theoretical survey literature, providing both a rigorous benchmark and methodological guidance for future research and industrial deployment of SSMs.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
Existing work lacks a theoretical explanation for how structured state space models (SSMs) support in-context learning (ICL) via gradient descent. Method: The authors explicitly construct a single-layer, gated SSM architecture with multiplicative input/output gating, capable of exactly simulating implicit linear and nonlinear model behavior under one- to multi-step gradient updates. Contribution/Results: This construction establishes a formal theoretical connection between SSMs and linear self-attention, identifying multiplicative gating as a critical inductive bias enabling large-model-like expressivity in recurrent architectures. Empirical validation confirms that randomly initialized models, after training, yield parameters closely matching analytical solutions; moreover, the proposed model successfully reproduces ICL capabilities on both linear and nonlinear regression tasks—demonstrating that gradient-based adaptation emerges intrinsically from the SSM’s structure and gating mechanism.
To address the inefficiency of global modeling and the superlinear computational growth in processing high-resolution images for vision tasks, this paper introduces VMamba—the first efficient state-space model family designed specifically for vision. The core innovation is a two-dimensional selective scanning (SS2D) mechanism that enables linear-complexity global contextual modeling over images via four-directional feature traversal, effectively bridging the fundamental gap between 1D sequential modeling and the non-sequential, 2D structure of images. Complemented by hardware-aware operator fusion and co-design optimizations at both architectural and implementation levels, VMamba achieves state-of-the-art or ViT/ConvNeXt–competitive accuracy on image classification, detection, and segmentation benchmarks. Crucially, its inference latency scales strictly linearly with input resolution, markedly improving efficiency for large-scale image processing.
This work addresses the limited interpretability of existing state space models regarding their long-range dependency mechanisms, particularly the unclear relationship between modeling capacity and architectural design in real-world tasks. Focusing on the S4D model, we present the first systematic analysis of its kernel behavior in the context of source code vulnerability detection. By integrating time-domain and frequency-domain analyses, we demonstrate that S4D can function as a low-pass, band-pass, or high-pass filter depending on its architectural configuration. This finding reveals that the model’s ability to capture long-range dependencies is profoundly influenced by its architecture, thereby offering both theoretical insights and concrete guidance for designing more effective state space models.
This paper addresses the fundamental question of why deep selective state space models (e.g., Mamba) efficiently model long-range dependencies. Method: It introduces rough path theory—novel in this context—to provide a rigorous mathematical foundation, modeling hidden states as low-dimensional projections of the input path signature and integrating selective state updates with input-controllable transitions to reveal their intrinsic capacity for capturing nonlinear token interactions across temporal scales. Contributions: (1) It establishes an expressivity upper bound for selective SSMs, proving their higher-order temporal modeling capability substantially surpasses that of conventional linear SSMs; (2) it unifies the explanation for the concurrent gains in accuracy and efficiency of Mamba-like architectures on continuous, long-sequence tasks (e.g., speech and video); (3) it provides a theoretically grounded, verifiable, and scalable framework—along with design principles—for next-generation structured state space models.
To address the insufficient spatial structure modeling and high computational cost of long-range dependency capture in 2D vision tasks, this paper proposes the Structure-Aware State Space Model (S3M). S3M introduces a novel structure-aware state fusion equation, unifying the theoretical frameworks of Mamba and linear attention. Its three-stage architecture integrates dilated convolutions, unidirectional scanning, and an observation equation to efficiently model pixel-wise neighborhood connectivity in a single scan. Unlike conventional state space models (SSMs), S3M explicitly encodes 2D spatial relationships without increasing scan path length. Extensive experiments demonstrate that S3M consistently outperforms existing SSM-based methods on image classification, object detection, and semantic segmentation—achieving superior efficiency and accuracy in long-range dependency modeling.
This work systematically compares how State Space Models (SSMs) and Transformers propagate contextual representations in long-sequence modeling. We propose the first unified analytical framework—integrating centered kernel alignment, stability metrics, probing experiments, and parameter randomization—to quantify inter-layer and inter-token information flow differences. Our analysis reveals that Transformers suffer from rapid representational homogenization (over-smoothing) of early tokens due to self-attention, whereas SSMs preserve representational diversity initially and converge gradually in deeper layers. Crucially, Transformer inductive bias arises primarily from architectural design, while SSM behavior is predominantly shaped by training dynamics. This study provides the first principled, interpretable characterization of fundamental representational divergence between these architectures, yielding actionable design principles and optimization guidelines for long-context modeling. (149 words)
Long-context language models suffer from low inference efficiency on consumer-grade hardware. Method: This work presents the first systematic, end-to-end performance evaluation of Transformer, State Space Model (SSM), and hybrid architectures on embedded and consumer GPUs (e.g., 24 GB VRAM), leveraging operator-level fine-grained analysis and real hardware benchmarking. Contribution/Results: We find SSMs significantly outperform Transformers for ultra-long sequences (up to 220K tokens), with performance crossover occurring at ~57K tokens and peak speedups of 4×. Moreover, SSMs support sequence lengths four times longer than Transformers on a 24 GB GPU. Crucially, we identify that over 55% of SSM latency stems from custom operators—revealing the primary bottleneck for edge deployment. This finding provides empirical grounding and a clear optimization pathway for efficient SSM implementation and hardware-software co-design on resource-constrained devices.
This work proposes NeuroSSM, an end-to-end multi-scale selective state space model that addresses the challenge of efficiently modeling the coexisting fast transient and slow large-scale dynamics in fMRI time series—a limitation of existing deep learning approaches that often rely on functional connectivity preprocessing. NeuroSSM uniquely integrates a multi-scale state space architecture with a parallel differential mechanism to directly process raw BOLD signals, enabling unified capture of multi-scale temporal dynamics. By doing so, it significantly enhances sensitivity to transient neural changes and achieves state-of-the-art performance across both clinical and non-clinical datasets. The method outperforms prevailing fMRI analysis techniques in both analytical accuracy and computational efficiency, offering a promising framework for direct, interpretable modeling of complex brain dynamics without intermediate preprocessing steps.
This study addresses the scalability limitations of Transformers in long-context modeling, where their O(N²) computational complexity becomes prohibitive. For the first time, it systematically evaluates the performance of the Mamba state space model against the LLaMA Transformer on real-world psychotherapy dialogue data across multi-scale context lengths (512–8192 tokens). The comparison quantifies the advantages of state space models along two dimensions: computational efficiency (memory footprint and inference latency) and representational efficiency (hidden state dynamics and attention patterns). Results demonstrate that Mamba substantially reduces computational overhead while preserving effective semantic representation under specific conditions, offering empirical evidence and practical guidance for model selection and deployment in long-context applications.
This work addresses the limitations of existing vision state space models, which rely on fixed-direction scanning strategies that often introduce redundancy and disrupt two-dimensional spatial dependencies. To overcome this, the authors propose MFil-Mamba, a novel architecture featuring a multi-filter scanning mechanism that dynamically captures context-aware spatial information. By adaptively weighting and fusing outputs from multiple scanning paths, the model effectively preserves complex spatial structures. The proposed method achieves state-of-the-art performance across multiple benchmarks: 83.2% top-1 accuracy on ImageNet-1K, 47.3% box AP and 42.7% mask AP on MS COCO for object detection and instance segmentation respectively, and 48.5% mIoU on ADE20K for semantic segmentation, consistently outperforming current best approaches.