A Comparative Analysis of Attention versus State-Space Models for In-Context Learning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified theoretical framework for comparing Transformers and state space models (SSMs) in in-context learning. To bridge this gap, we propose the "belief geometry" analytical framework, which abstracts evidence assembly, belief maintenance, and addressing capabilities. Our methodology employs cumulative Bayesian regret metrics and generalized contextual linear regression modeling, complemented by empirical validation using LLaMA and Mamba-2. This work establishes the first general theoretical framework demonstrating the optimality of SSMs in memory maintenance alongside the exponential width advantage of attention mechanisms in content-based addressing. Crucially, these insights are successfully extended to nonlinear settings, providing clear theoretical guidance for architecture selection in sequence learning tasks.
📝 Abstract
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.
Problem

Research questions and friction points this paper is trying to address.

State-Space Models
Attention Mechanisms
In-Context Learning
Transformers
Belief Geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

Belief Geometry
State-Space Models
Attention Mechanisms
In-Context Learning
Bayes Regret
🔎 Similar Papers
No similar papers found.