Adapting Linear-Time Architectures for Tabular In-Context Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that linear architectures struggle to generalize beyond pre-trained sequence lengths in tabular in-context learning due to recurrent state drift. To overcome this, we optimize the DeltaNet architecture by uncovering the theoretical origins of state instability in causal models and introducing a time-dependent decay scheduling strategy to stabilize recurrent states. Furthermore, a non-causal readout mechanism is incorporated to enhance long-sequence modeling capabilities. The proposed approach effectively transcends pre-trained length constraints and significantly improves length extrapolation robustness. Evaluated on the OpenML-CC18 and TabArena benchmarks, our method achieves performance comparable to Softmax attention baselines, demonstrating that linear attention provides a reliable solution for efficient long-range tabular reasoning.
📝 Abstract
Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large datasets. Existing linear-time alternatives, however, are mostly causal, and their potential for tabular in-context learning (ICL) remains underexplored. To address this, we (1) revisit causal training setups, (2) compare linear sequence mixers, and (3) investigate their ICL generalisation beyond the pretraining context length. First, we show that the best training setup for causal models resembles next-token prediction. Then, perhaps surprisingly, the most promising linear sequence mixer is causal: DeltaNet outperforms even non-causal linear attention. However, it degrades beyond $2$-$4\times$ the pretraining context length, and existing mitigation strategies such as bidirectionality defer the problem at best. A hidden-state oracle shows that this is not a capacity problem. Instead, our analysis points to an instability in the recurrent state, which drifts in deeper layers of causal models. Since DeltaNet's learned write rates overfit to the pretraining regime, we modulate them with a time-dependent decay schedule intervention to stabilise length generalisation. Finally, re-introducing non-causality by reading out from the final state allows us to closely match a controlled softmax attention baseline on OpenML-CC18 and TabArena.
Problem

Research questions and friction points this paper is trying to address.

Tabular In-Context Learning
Linear-Time Architectures
Causal Models
Length Generalization
Recurrent State Instability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tabular In-Context Learning
Linear-Time Architectures
DeltaNet
Length Generalization
Recurrent State Stability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
David Schnurr
ETH Zürich
F
Felix Sarnthein
ELLIS Institute Tübingen, MPI-IS
T
Thomas Hofmann
ETH Zürich, ETH AI Center
Imanol Schlag
Imanol Schlag
ETH AI Center
Responsible AILarge Language ModelsAssociative RNNs / DeltaNet