dual-stream causal attention

Designs, implements, and analyzes transformer-style attention mechanisms that maintain two separate autoregressive (causal) streams and the rules for how those streams attend to themselves and to each other. This includes creating causal mask patterns, cross-stream coordination logic, and efficient computation/implementation details for two-stream or dual-stream causal transformers so the model enforces per-stream causality while supporting controlled inter-stream information flow.

dual-streamcausalattention

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge that standard neural models struggle to reliably reproduce the optimization trajectories of causal discovery algorithms. To this end, it proposes a fixed-weight linear attention Transformer that precisely executes continuous causal graph updates by retaining Lagrange multipliers. Theoretically, we prove that multiplier retention is essential for exact execution, establish conditions under which rounding errors remain bounded as network depth increases, and decouple algorithmic execution from causal recovery capabilities. Empirically, the proposed model achieves floating-point precision consistency on synthetic data and seven benchmark networks while inheriting the performance of reference solvers. These results validate the feasibility of employing Transformers to exactly execute causal discovery algorithms.

Algorithm ExecutionCausal DiscoveryCausal Structure Learning

This study addresses the challenge of accurately discovering time-lagged causal structures from multivariate time series without explicit causal constraints. The work proposes leveraging the autoregressive Transformer’s inherent sensitivity—specifically, the gradient of its predictions with respect to historical inputs—as a natural encoder of causal relationships, from which a causal graph is extracted via gradient attribution aggregation. It is rigorously demonstrated for the first time that standard Transformers possess intrinsic causal learning capabilities without requiring additional causal objectives or architectural modifications. Notably, causal discovery accuracy improves significantly as data heterogeneity increases. The method substantially outperforms existing causal discovery algorithms under challenging conditions, including nonlinear dynamics, long-range dependencies, and nonstationarity, with particularly pronounced gains in highly heterogeneous datasets.

autoregressivecausal discoverycausal graph

This work addresses the lack of a unified theoretical foundation in existing attention mask designs. It establishes, for the first time, a formal connection between attention masks and partially ordered structures, proving that information flow in sufficiently deep multi-layer Transformers converges to a Hasse diagram. The mask design problem is thereby reformulated as finding the minimal common supergraph of such Hasse diagrams, yielding a general framework that derives attention masks directly from task families. Leveraging this framework, the authors propose two novel mechanisms—Block Two-Stream Attention and Butterfly Attention—and derive block-wise causal masks and fully supervised bidirectional masks that guarantee consistency between training and inference. Empirical results validate both the effectiveness and generality of the proposed approach.

attention masksHasse diagramsinformation flow

This work addresses the “lost-in-the-middle” phenomenon in Transformer decoders—where information located in the middle of long prompts is poorly retrieved—by providing, for the first time, a rigorous theoretical explanation from a dynamical systems perspective. By modeling causal self-attention as a non-commutative interacting particle system and leveraging cumulant expansions under triangular causal structure together with Glauber calculus, the authors establish a mean-field limit and analytically solve the associated correlation equations. Under the assumption of i.i.d. uniformly distributed inputs, they rigorously prove that token retrieval performance exhibits a U-shaped dependence on source position, thereby quantitatively elucidating the origins of primacy and recency effects alongside the characteristic performance dip in the middle of the sequence.

causal dynamicslost-in-the-middleself-attention

Towards Understanding the Universality of Transformers for Next-Token Prediction

Oct 03, 2024
ME
Michael E. Sander
🏛️ Ecole normale superieure | CNRS

This work investigates the universal mechanism by which Transformers implicitly learn temporal mappings—such as linear and periodic patterns—in autoregressive next-token prediction, with a focus on how self-attention encodes causal sequence structure. We propose *causal kernel descent*, a novel theoretical framework that, for the first time, constructsively proves that causal Transformers can implicitly implement an online Kaczmarz-type algorithm via self-attention to approximate a dynamic context function (f) in a Hilbert space. We rigorously construct a Transformer architecture capable of exactly learning (f), and validate its efficacy through controlled sequence experiments under linear, exponential, and softmax attention. This work establishes the first provably correct, constructively universal theoretical framework for understanding how large language models intrinsically perform in-context reasoning—offering a new paradigm for analyzing their implicit inference capabilities.

Developing a causal kernel descent method for in-context learning.Exploring causal Transformers' capacity for autoregressive sequence prediction.Understanding Transformers' universal next-token prediction ability.

Latest Papers

What's happening recently
View more

This study investigates the expressive power of causal-masked Transformers with finite-precision arithmetic in solving decision problems over arbitrarily long input sequences. By integrating algebraic formalism with finite-precision numerical semantics and semigroup theory, the authors develop a compositional framework centered on memory states to analyze representational capacity. They establish, for the first time, a precise correspondence between four classes of attention mechanisms and specific semigroup varieties—namely, aperiodic, R-trivial, locally R-trivial, and aperiodic semigroups—and prove that, under the free wiring assumption, the expressiveness boundaries of these classes are all tight. This work offers a novel algebraic perspective for understanding the theoretical capabilities of Transformers under realistic computational constraints.

attention mechanismscausally masked transformersdecision problems

This work addresses the lack of a formal logical characterization for encoder-decoder Transformers employing floating-point soft attention. The authors propose a novel temporal logic that integrates a counting global modality tailored to encoder inputs and a past modality designed for decoder inputs. Coupled with a distributed automaton model, this framework constitutes the first logical characterization suitable for practical floating-point soft attention settings. It naturally accommodates common architectural variants such as masking, establishes an equivalence between Transformers and distributed automata, and extends seamlessly to autoregressive scenarios. The approach demonstrates both robustness to architectural variations and strong expressive power, thereby providing a foundational formalism for reasoning about realistic Transformer models.

cross-attentionencoder-decoder transformerslogical characterization

This work addresses the limitation of conventional Transformers, which employ fixed additive residual connections along the depth axis and thus lack adaptive inter-layer information aggregation. The authors propose a dual perspective on residual flows, conceptualizing the decoder as an evolving system across two ordered dimensions—sequence positions and layer depth—and establish, for the first time, an operator-level equivalence between causal residual attention along the depth axis and short sliding window attention (ShortSWA) along the sequence axis. This duality provides a unified interpretation of methods such as ELC-BERT, DenseFormer, and Vertical Attention, while distinguishing operator-level duality from system-level asymmetry. Building on this insight, the study integrates Deep Delta Learning (DDL) with ShortSWA to demonstrate that, in large-scale autoregressive models, ShortSWA aligns better with hardware characteristics, whereas DDL more effectively optimizes residual pathways, thereby offering clear principles for architectural design.

adaptive mixingcross-layer informationdepth-wise attention

This work addresses the challenge of moving beyond correlational analyses to establish causal relationships in identifying truly influential features within Transformer models. It proposes a five-stage causal analysis framework—comprising probe design, feature extraction, causal validation, robustness testing, and deployment integration—to systematically evaluate features in GPT-2 small on the Indirect Object Identification (IOI) task. By integrating activation patching, sparse autoencoders, feature ablation, and an NLA-inspired variance attribution method, the study reveals a negative correlation between feature selectivity and causal efficacy, and uncovers a substantial gap between detection robustness and causal robustness. While replicating the IOI circuit, it finds that only a subset of 15 highly selective features exhibit genuine causal effects, and model performance degrades significantly under distributional shift. A cost-aware monitoring strategy achieves 99.1% resource savings.

causal feature analysiscorrelation vs. causationfeature causality

This study investigates the distinction between the causal use and mere decodability of internal representations in Transformers when performing hierarchical tasks. By integrating representation probing, attention masking, and residual stream subspace ablation on Dyck languages and templated natural language tasks, the work provides the first clear separation between the decodability of hierarchical information and its causal role in model computation. The findings reveal that although hierarchical signals are widely decodable across representations, the model relies selectively on specific mechanisms—such as attention to stack-top positions—to handle long-range dependencies. Moreover, ablating low-dimensional residual subspaces has negligible impact on performance, demonstrating that decodability does not imply causal utilization. This work thus uncovers a critical gap between the readability of internal representations and the actual mechanisms driving model reasoning.

causal usedecodabilityDyck language

Hot Scholars

WX

Weidi Xie

Shanghai Jiao Tong University | VGG, University of Oxford
Computer VisionAI for HealthcareAI for Science
WZ

Wenzhao Zheng

EECS, University of California, Berkeley
Large ModelsEmbodied AgentsAutonomous Driving
SD

Shangzhe Di

Shanghai Jiao Tong University
Video UnderstandingMultimodal LearningComputer Vision
YY

Yibin Yan

PhD, Shanghai Jiao Tong University
Computer VisionMultimodal PerceptionVideo Understanding
JX

Jilan Xu

Fudan University
Computer VisionMultimodalMedical Image Analysis