Causal and Interpretable Structures in LLM Compositional Tasks

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how large language models represent and process relational information in compositional tasks. By analyzing the geometric structure and causal mechanisms of activations across Transformer layers, combined with prompt ensembling and causal intervention techniques, we conduct cross-model experiments on mainstream architectures such as Llama. Our findings reveal a hierarchical progression in recurrent conceptual reasoning, uncovering an incremental organizational mechanism wherein intermediate layers exploit binary relations while later layers leverage ternary relations, alongside the identification of causally inert structures. Furthermore, we demonstrate that constraining models to rely exclusively on causally relevant joint representations significantly improves next-token prediction accuracy.
📝 Abstract
Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Compositional Tasks
Causal Representation
Relational Information
Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Representation
Compositional Generalization
Mechanistic Interpretability
Joint Representations
Large Language Models