non-causal ssm design

Designs, implements, and evaluates non-causal state-space model (SSM) architectures and their parameterizations that process tokens without a causal (left-to-right) constraint, enabling full-context or permutation-invariant token interactions in a single pass. This work produces SSM variants and inference schemes that remove directional scanning bias, achieve low single-sample inference latency, and preserve spatial symmetry or orientation-invariance when operating on structured inputs.

non-causalssmdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

State Space Model Programming in Turing.jl

May 29, 2025
TH
Tim Hargreaves
🏛️ University of Cambridge | Federal Reserve Board of Governors

Existing state-space model (SSM) frameworks lack composability and scalability, hindering rapid model experimentation and integration of advanced inference techniques. Method: We introduce the first modular SSM programming framework built on the Julia ecosystem, decoupling model structure from inference algorithms to enable unified modeling and inference across linear Gaussian, nonlinear, and non-Gaussian systems. The framework integrates Turing.jl, SSMProblems.jl, and GeneralisedFilters.jl; introduces a novel dynamic scheduling mechanism for hybrid filtering (e.g., Kalman + particle filters); and incorporates CUDA-accelerated GPU execution with memory optimization. Contribution/Results: Experiments demonstrate substantially reduced development complexity—enabling real-time inference on million-scale time-series data—while improving code reuse by over 3× and achieving an 8.2× throughput gain for GPU-based inference over CPU counterparts.

Difficulty in applying advanced inference techniques to SSMsLack of composable, scalable frameworks for state space modelsNeed for efficient, modular SSM implementations in time-series analysis

SaFARi: State-Space Models for Frame-Agnostic Representation

May 13, 2025
HB
Hossein Babaei
🏛️ Rice University

Existing state space models (SSMs) heavily rely on restrictive polynomial bases—such as HiPPO—limiting their adaptability to arbitrary signal representation frameworks and thereby impairing flexibility and robustness in long-range dependency modeling. To address this, we propose a frame-agnostic, universal SSM construction paradigm that rigorously generalizes SSMs to arbitrary orthonormal bases—including Fourier and wavelet bases—as well as general Parseval frames. Methodologically, our approach unifies functional space analysis with structured state-space modeling by integrating generalized orthogonal projection theory with continuous-time system discretization techniques. This framework subsumes HiPPO as a special case while enabling infinitely many novel SSM instantiations. Crucially, it preserves theoretical soundness while substantially enhancing robustness to sequence length variation and expanding modeling diversity across signal domains.

Enabling diverse frame-agnostic representations in SSMsExtending HiPPO framework to infinite basis optionsGeneralizing SSM construction beyond polynomial bases

Theoretical Foundations of Deep Selective State-Space Models

Feb 29, 2024
NM
Nicola Muca Cirone
🏛️ Imperial College London | MPI for Intelligent Systems | University of Oxford

This paper addresses the fundamental question of why deep selective state space models (e.g., Mamba) efficiently model long-range dependencies. Method: It introduces rough path theory—novel in this context—to provide a rigorous mathematical foundation, modeling hidden states as low-dimensional projections of the input path signature and integrating selective state updates with input-controllable transitions to reveal their intrinsic capacity for capturing nonlinear token interactions across temporal scales. Contributions: (1) It establishes an expressivity upper bound for selective SSMs, proving their higher-order temporal modeling capability substantially surpasses that of conventional linear SSMs; (2) it unifies the explanation for the concurrent gains in accuracy and efficiency of Mamba-like architectures on continuous, long-sequence tasks (e.g., speech and video); (3) it provides a theoretically grounded, verifiable, and scalable framework—along with design principles—for next-generation structured state space models.

Continuous Large-scale DataDeep Selective State Space ModelsHidden State Influence

Language Models as Causal Effect Generators

Nov 12, 2024
LE
Lucius E.J. Bynum
🏛️ New York University | Prescient Design | Genentech

Existing benchmarks for evaluating causal inference methods and auditing implicit causal reasoning in large language models (LLMs) lack controllability and the capacity to generate counterfactual data. Method: We propose Sequence-Driven Structural Causal Models (SD-SCMs), a novel framework that treats LLMs as structural equation providers—integrated with user-specified directed acyclic graphs (DAGs)—to automatically generate observational, interventional, and individual-level counterfactual datasets. Contribution/Results: (1) SD-SCMs enable interpretable SCM construction without manual specification of functional forms; (2) we introduce a new causal benchmark comprising thousands of heterogeneous datasets, enabling systematic evaluation of effect estimators’ robustness under varying confounding conditions; (3) we provide an auditable, attributable mechanism for detecting implicit causal effects embedded within LLMs. This framework bridges the gap between causal reasoning evaluation and generative model introspection, supporting rigorous, scalable, and reproducible causal assessment.

Auditing language models for desirable and undesirable causal effectsCreating benchmarks to test treatment effect estimation methodsProposing a framework for causal models with language-defined mechanisms

This work investigates the fundamental expressive capacity of linear state-space models (SSMs) for language modeling, clarifying their theoretical modeling boundaries relative to Transformers and classical RNNs. Method: Leveraging formal language theory and automata theory, we formally characterize SSM expressivity—proving for the first time that linear SSMs can exactly recognize star-free languages and optimally model bounded hierarchical structures in memory. We identify a critical expressivity bottleneck in contemporary SSM designs arising from the absence of nonlinearity in state updates. Contribution/Results: Our analysis reveals that SSMs and Transformers possess complementary—not substitutive—capabilities. Empirical evaluation on the Mamba architecture demonstrates substantial gains over Transformers on star-free language tasks and superior memory efficiency in hierarchical structure modeling. These findings provide both theoretical foundations and practical guidance for designing next-generation efficient large language model architectures.

Current SSMs have design limits affecting their expressivenessSSMs handle star-free state tracking better than transformersSSMs' expressive power compared to transformers and RNNs

Latest Papers

What's happening recently
View more

This work addresses the high computational and memory demands of Structured State Space Models (e.g., S4/S4D), which, despite their effectiveness in modeling long-range dependencies, hinder deployment on resource-constrained devices. The study presents the first systematic exploration of operator-level structured pruning tailored for S4/S4D, introducing an incremental mask generation strategy, a joint accuracy-latency monitoring mechanism, and a unified training-evaluation framework. By alternately applying structured pruning and fine-tuning, the method achieves efficient model compression. Experiments demonstrate that up to 70% of operators can be pruned across multiple benchmarks with negligible performance degradation, substantially reducing inference latency and offering a practical pathway toward lightweight deployment of state space models.

inference latencymodel efficiencyoperator-level pruning

This work addresses the limitation of existing probing methods, which struggle to identify reusable internal causal interfaces in language models that support diverse future computations due to their reliance solely on current outputs. The authors propose a label-free framework for discovering such causal interfaces by introducing a “forked futures” mechanism: after a shared prefix, multiple divergent future trajectories are sampled to construct a causal quotient space based on comparisons of response distributions. Four interface types—including Shared—are defined and competitively evaluated using a preordered causal description length, with structural selection guided by a fidelity constraint on future signatures. Experiments demonstrate that the method reduces description length by 0.216 and 0.294 nats on Qwen2.5-1.5B and Llama-3-8B, respectively; it successfully recovers 14 out of 16 model architectures in blind tests and achieves an API alignment path mediation effect of 0.749.

causal interfacesforked futureshidden APIs

This study addresses the structural errors in language models arising from suboptimal causality assumptions and the absence of systemic behavior by proposing the SBD framework. This framework is the first to incorporate systemic behavior as an irreducible component into Bayesian features, revealing the "causality tax" phenomenon. It further constructs a non-causal variational family, GSH, to replace conventional sequential dependency chains, and optimizes foundation models by integrating ELBO theory, implicit measurement, NTK evaluation, and divide-and-conquer strategies. Experimental results demonstrate that GSH achieves a signal-to-noise ratio improvement exceeding 7 dB and enhances multi-scale fitting capability by 20%, significantly outperforming existing causal models.

Causality TaxData DistributionLanguage Models

This study demonstrates that even in functionally and performance-wise correct distributed AI inference systems, microsecond-level clock skew among nodes can induce observable causality violations. By injecting controlled clock offsets into a multi-node inference pipeline built on Kafka and ZeroMQ, the authors reveal—for the first time—the high sensitivity of such causal anomalies to temporal synchronization: as little as 5 ms of offset suffices to produce noticeable violations, with their manifestation dynamically evolving alongside relative clock drift. Crucially, while system throughput and output correctness remain unaffected, observability degrades significantly, underscoring the necessity of treating time as a first-class concern in the design and operation of distributed AI systems.

causality violationsclock skewdistributed AI inference

This work addresses a critical limitation in existing selective state space models—such as Mamba—which lose the physical interpretability of time steps by treating them as arbitrary input-dependent functions, thereby struggling with irregularly sampled time series. While continuous-time models like S5 preserve temporal semantics, they are constrained by linear time-invariant dynamics and lack per-token expressivity. To overcome these issues, we propose TIDES, the first selective state space model that decouples input dependence from time steps and instead introduces an input-dependent diagonal state matrix. This design retains the physical meaning of time steps while achieving high representational capacity. TIDES natively supports irregular sampling, attains the best average rank on UEA classification and Physiome-ODE regression benchmarks, and demonstrates strong out-of-distribution extrapolation on a newly introduced Fading Flash task involving unseen time intervals.

continuous-time modelingirregular time seriesper-token expressivity

Hot Scholars

YB

Yun-Bo Zhao

University of Science and Technology of China
Human-Machine SystemsSmart ManufacturingNetworked Control Systems
SB

Samir Bhatt

Professor of Machine Learning and Public Health University of Copenhagen
Public HealthGeneticsInfectious DiseasesMachine Learning
HB

Harrison Bo Hua Zhu

Assistant Professor, University of Copenhagen
EpidemiologyPhylogeneticsInfectious DiseasesProbabilistic Machine Learning
JH

Jianshu Hu

Shanghai Jiao Tong University
Reinforcement LearningRobotics
YL

Yingzhen Li

Imperial College London
Artificial IntelligenceMachine LearningStatistics