ssm-based depth modeling

Designs and implements state‑space model layers or modules that aggregate and propagate long‑range dependencies along the depth (slice/volume) axis of data, providing efficient depth‑wise context modeling. Integrates these state‑space depth modules into hybrid 2D/3D neural architectures and analyzes their effects on segmentation stability, variance, and overall performance.

ssm-baseddepthmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

State-Space Modeling in Long Sequence Processing: A Survey on Recurrence in the Transformer Era

Jun 13, 2024
MT
Matteo Tiezzi
🏛️ IIT | University of Siena | IMT

Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.

Addressing limitations of Transformers with state-space modelsExploring efficient online learning beyond backpropagation through timeSurveying recurrent models for long sequence processing

Must-Read Papers

Most classic and influential ideas
View more

State-space models can learn in-context by gradient descent

Oct 15, 2024
NM
Neeraj Mohan Sushma
🏛️ Ruhr Universität Bochum | Birla Institute of Technology and Science | Royal Holloway, University of London

Existing work lacks a theoretical explanation for how structured state space models (SSMs) support in-context learning (ICL) via gradient descent. Method: The authors explicitly construct a single-layer, gated SSM architecture with multiplicative input/output gating, capable of exactly simulating implicit linear and nonlinear model behavior under one- to multi-step gradient updates. Contribution/Results: This construction establishes a formal theoretical connection between SSMs and linear self-attention, identifying multiplicative gating as a critical inductive bias enabling large-model-like expressivity in recurrent architectures. Empirical validation confirms that randomly initialized models, after training, yield parameters closely matching analytical solutions; moreover, the proposed model successfully reproduces ICL capabilities on both linear and nonlinear regression tasks—demonstrating that gradient-based adaptation emerges intrinsically from the SSM’s structure and gating mechanism.

Exploring SSMs' relationship with linear self-attentionProving SSMs can perform gradient-based learningUnderstanding in-context learning in state-space models

VMamba: Visual State Space Model

Jan 18, 2024
YL
Yue Liu
🏛️ UCAS | Huawei Inc. | Pengcheng Lab

To address the inefficiency of global modeling and the superlinear computational growth in processing high-resolution images for vision tasks, this paper introduces VMamba—the first efficient state-space model family designed specifically for vision. The core innovation is a two-dimensional selective scanning (SS2D) mechanism that enables linear-complexity global contextual modeling over images via four-directional feature traversal, effectively bridging the fundamental gap between 1D sequential modeling and the non-sequential, 2D structure of images. Complemented by hardware-aware operator fusion and co-design optimizations at both architectural and implementation levels, VMamba achieves state-of-the-art or ViT/ConvNeXt–competitive accuracy on image classification, detection, and segmentation benchmarks. Crucially, its inference latency scales strictly linearly with input resolution, markedly improving efficiency for large-scale image processing.

Computational EfficiencyImage ProcessingLarge Images

This work addresses the limited interpretability of existing state space models regarding their long-range dependency mechanisms, particularly the unclear relationship between modeling capacity and architectural design in real-world tasks. Focusing on the S4D model, we present the first systematic analysis of its kernel behavior in the context of source code vulnerability detection. By integrating time-domain and frequency-domain analyses, we demonstrate that S4D can function as a low-pass, band-pass, or high-pass filter depending on its architectural configuration. This finding reveals that the model’s ability to capture long-range dependencies is profoundly influenced by its architecture, thereby offering both theoretical insights and concrete guidance for designing more effective state space models.

interpretabilitykernel analysislong-range dependency

Theoretical Foundations of Deep Selective State-Space Models

Feb 29, 2024
NM
Nicola Muca Cirone
🏛️ Imperial College London | MPI for Intelligent Systems | University of Oxford

This paper addresses the fundamental question of why deep selective state space models (e.g., Mamba) efficiently model long-range dependencies. Method: It introduces rough path theory—novel in this context—to provide a rigorous mathematical foundation, modeling hidden states as low-dimensional projections of the input path signature and integrating selective state updates with input-controllable transitions to reveal their intrinsic capacity for capturing nonlinear token interactions across temporal scales. Contributions: (1) It establishes an expressivity upper bound for selective SSMs, proving their higher-order temporal modeling capability substantially surpasses that of conventional linear SSMs; (2) it unifies the explanation for the concurrent gains in accuracy and efficiency of Mamba-like architectures on continuous, long-sequence tasks (e.g., speech and video); (3) it provides a theoretically grounded, verifiable, and scalable framework—along with design principles—for next-generation structured state space models.

Continuous Large-scale DataDeep Selective State Space ModelsHidden State Influence

Spatial-Mamba: Effective Visual State Space Models via Structure-Aware State Fusion

Oct 19, 2024
CX
Chaodong Xiao
🏛️ The Hong Kong Polytechnic University | OPPO Research Institute | Harvard Medical School | Xi’an Jiaotong University

To address the insufficient spatial structure modeling and high computational cost of long-range dependency capture in 2D vision tasks, this paper proposes the Structure-Aware State Space Model (S3M). S3M introduces a novel structure-aware state fusion equation, unifying the theoretical frameworks of Mamba and linear attention. Its three-stage architecture integrates dilated convolutions, unidirectional scanning, and an observation equation to efficiently model pixel-wise neighborhood connectivity in a single scan. Unlike conventional state space models (SSMs), S3M explicitly encodes 2D spatial relationships without increasing scan path length. Extensive experiments demonstrate that S3M consistently outperforms existing SSM-based methods on image classification, object detection, and semantic segmentation—achieving superior efficiency and accuracy in long-range dependency modeling.

Captures complex image spatial structuresEnhances 2D vision task performanceReduces computational cost in SSMs

Latest Papers

What's happening recently
View more

A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures

Oct 08, 2025
NM
Nhat M. Hoang
🏛️ Nanyang Technological University | National University of Singapore

This work systematically compares how State Space Models (SSMs) and Transformers propagate contextual representations in long-sequence modeling. We propose the first unified analytical framework—integrating centered kernel alignment, stability metrics, probing experiments, and parameter randomization—to quantify inter-layer and inter-token information flow differences. Our analysis reveals that Transformers suffer from rapid representational homogenization (over-smoothing) of early tokens due to self-attention, whereas SSMs preserve representational diversity initially and converge gradually in deeper layers. Crucially, Transformer inductive bias arises primarily from architectural design, while SSM behavior is predominantly shaped by training dynamics. This study provides the first principled, interpretable characterization of fundamental representational divergence between these architectures, yielding actionable design principles and optimization guidelines for long-context modeling. (149 words)

Analyzing contextual representation flow in state-space and transformer architecturesComparing token homogenization patterns between SSMs and transformer modelsIdentifying architectural versus training causes of representation oversmoothing

Long-context language models suffer from low inference efficiency on consumer-grade hardware. Method: This work presents the first systematic, end-to-end performance evaluation of Transformer, State Space Model (SSM), and hybrid architectures on embedded and consumer GPUs (e.g., 24 GB VRAM), leveraging operator-level fine-grained analysis and real hardware benchmarking. Contribution/Results: We find SSMs significantly outperform Transformers for ultra-long sequences (up to 220K tokens), with performance crossover occurring at ~57K tokens and peak speedups of 4×. Moreover, SSMs support sequence lengths four times longer than Transformers on a 24 GB GPU. Crucially, we identify that over 55% of SSM latency stems from custom operators—revealing the primary bottleneck for edge deployment. This finding provides empirical grounding and a clear optimization pathway for efficient SSM implementation and hardware-software co-design on resource-constrained devices.

Comparing performance of SSMs and Transformers on consumer hardwareEvaluating SSM and hybrid models for long-context processing efficiencyIdentifying hardware-aware optimizations for edge device applications

This work proposes NeuroSSM, an end-to-end multi-scale selective state space model that addresses the challenge of efficiently modeling the coexisting fast transient and slow large-scale dynamics in fMRI time series—a limitation of existing deep learning approaches that often rely on functional connectivity preprocessing. NeuroSSM uniquely integrates a multi-scale state space architecture with a parallel differential mechanism to directly process raw BOLD signals, enabling unified capture of multi-scale temporal dynamics. By doing so, it significantly enhances sensitivity to transient neural changes and achieves state-of-the-art performance across both clinical and non-clinical datasets. The method outperforms prevailing fMRI analysis techniques in both analytical accuracy and computational efficiency, offering a promising framework for direct, interpretable modeling of complex brain dynamics without intermediate preprocessing steps.

BOLD signalsfMRI analysislong-range dependencies

This study addresses the scalability limitations of Transformers in long-context modeling, where their O(N²) computational complexity becomes prohibitive. For the first time, it systematically evaluates the performance of the Mamba state space model against the LLaMA Transformer on real-world psychotherapy dialogue data across multi-scale context lengths (512–8192 tokens). The comparison quantifies the advantages of state space models along two dimensions: computational efficiency (memory footprint and inference latency) and representational efficiency (hidden state dynamics and attention patterns). Results demonstrate that Mamba substantially reduces computational overhead while preserving effective semantic representation under specific conditions, offering empirical evidence and practical guidance for model selection and deployment in long-context applications.

computational efficiencylong-context sequencesrepresentational efficiency

This work addresses the limitations of existing vision state space models, which rely on fixed-direction scanning strategies that often introduce redundancy and disrupt two-dimensional spatial dependencies. To overcome this, the authors propose MFil-Mamba, a novel architecture featuring a multi-filter scanning mechanism that dynamically captures context-aware spatial information. By adaptively weighting and fusing outputs from multiple scanning paths, the model effectively preserves complex spatial structures. The proposed method achieves state-of-the-art performance across multiple benchmarks: 83.2% top-1 accuracy on ImageNet-1K, 47.3% box AP and 42.7% mask AP on MS COCO for object detection and instance segmentation respectively, and 48.5% mIoU on ADE20K for semantic segmentation, consistently outperforming current best approaches.

2D Spatial DependenciesComputer VisionSequential Modeling

Hot Scholars

RB

Rebekka Burkholz

CISPA Helmholtz Center for Information Security
Machine LearningDeep Learning EfficiencyComplex NetworksCascades
TF

Tingxiang Fan

The University of Hong Kong
RoboticsAutonomous DrivingMachine Learning
CZ

Chao Zhou

postdoc@CISPA; PhD@UCL
machine learning
XL

Xiangpeng Li

Chongqing University
Visual Question AnsweringCaptioningLLM
TJ

Tom Jacobs

PhD student, CISPA Helmholtz Center for Information Security
Training DynamicsDeep LearningOptimization