Institution profile

Goodfire AI

Industry researcheurope · gb
Official website
Research library38linked papers
Opportunities0open roles
Selected work

Representative Papers

Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics

Nov 06, 2025

This work investigates whether large language models (LLMs) implicitly represent unselected reasoning paths—i.e., “untraversed chains of thought”—during text generation. Method: We propose an uncertainty quantification framework based on dynamic analysis of hidden states: (i) decoding intermediate-layer activations to predict future token distributions, and (ii) conducting activation intervention experiments to assess hidden-state sensitivity to multi-path competition. Contribution/Results: We demonstrate that hidden states not only encode the current token decision but also explicitly embed the geometric structure of alternative reasoning paths in latent space. Crucially, the degree of uncertainty encoded in these states strongly correlates with model controllability—i.e., the ease with which generation can be steered toward desired outputs. This study provides the first empirical evidence for LLMs possessing “path awareness” and establishes a principled, interpretable linkage among hidden-state representations, predictive uncertainty, and generation controllability.

1 citationsRead paper

Inference and learning in sparse autoencoders as natural gradient flow

Oct 05, 2026

This study addresses the challenge of reliably recovering interpretable features using sparse autoencoders (SAEs) under feature superposition and low-frequency activation. To this end, we propose BeFOND, a model that unifies inference and dictionary learning into a natural gradient flow, enabling encoder-free sparse coding. Specifically, BeFOND eliminates feature interference through recurrent inference and compensates for the slow learning of rare features via Fisher information matrix preconditioning. Experimental results demonstrate that our approach significantly enhances both dictionary recovery and rare feature detection. Notably, BeFOND outperforms pretrained SAEs with substantially less data, and its performance continues to improve consistently as dictionary width scales up.

0 citationsRead paper

The Independence Prior of SAEs Fragments Visual Concepts

Oct 02, 2026

This study addresses the fragmentation of visual concepts in existing sparse autoencoders (SAEs) caused by neglecting image spatial dependencies. To this end, we propose the Markov random field linear representation hypothesis, which incorporates spatial dependency priors to refine the conventional independence assumption. Building upon this, we design Spatial-SAE as an amortized maximum a posteriori (MAP) estimator to recover spatial correlations. By integrating Markov random field modeling with SAEs, our approach achieves more precise disentanglement and interpretation of visual concepts. Experimental results demonstrate that Spatial-SAE attains a 96% success rate on synthetic concept recovery tasks and significantly enhances the interpretability of DINOv2 activations.

0 citationsRead paper

When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task

Sep 29, 2026

This study investigates how large language models leverage representational geometry for computation in numerical comparison tasks. Through a causal geometric analysis of the Qwen model, this work reveals that numerical comparison is achieved via the coordinated interplay of attention mechanisms, residual connections, and MLP neurons, which collectively enable local interval comparison and maximum-value localization. The findings demonstrate that the model primarily relies on linear representations rather than manifold operations for decision-making, while theoretically establishing the coexistence of the manifold hypothesis and linear representational structures. By elucidating a multi-number decision mechanism grounded in linear superposition and local comparison, this research challenges the prevailing overreliance on curvature geometry within mechanistic interpretability, offering novel perspectives for understanding internal model computations.

0 citationsRead paper

Causal and Interpretable Structures in LLM Compositional Tasks

Sep 28, 2026

This study investigates how large language models represent and process relational information in compositional tasks. By analyzing the geometric structure and causal mechanisms of activations across Transformer layers, combined with prompt ensembling and causal intervention techniques, we conduct cross-model experiments on mainstream architectures such as Llama. Our findings reveal a hierarchical progression in recurrent conceptual reasoning, uncovering an incremental organizational mechanism wherein intermediate layers exploit binary relations while later layers leverage ternary relations, alongside the identification of causally inert structures. Furthermore, we demonstrate that constraining models to rely exclusively on causally relevant joint representations significantly improves next-token prediction accuracy.

0 citationsRead paper
Recent publications

Latest Papers

Inference and learning in sparse autoencoders as natural gradient flow

Oct 05, 2026

This study addresses the challenge of reliably recovering interpretable features using sparse autoencoders (SAEs) under feature superposition and low-frequency activation. To this end, we propose BeFOND, a model that unifies inference and dictionary learning into a natural gradient flow, enabling encoder-free sparse coding. Specifically, BeFOND eliminates feature interference through recurrent inference and compensates for the slow learning of rare features via Fisher information matrix preconditioning. Experimental results demonstrate that our approach significantly enhances both dictionary recovery and rare feature detection. Notably, BeFOND outperforms pretrained SAEs with substantially less data, and its performance continues to improve consistently as dictionary width scales up.

0 citationsRead paper

The Independence Prior of SAEs Fragments Visual Concepts

Oct 02, 2026

This study addresses the fragmentation of visual concepts in existing sparse autoencoders (SAEs) caused by neglecting image spatial dependencies. To this end, we propose the Markov random field linear representation hypothesis, which incorporates spatial dependency priors to refine the conventional independence assumption. Building upon this, we design Spatial-SAE as an amortized maximum a posteriori (MAP) estimator to recover spatial correlations. By integrating Markov random field modeling with SAEs, our approach achieves more precise disentanglement and interpretation of visual concepts. Experimental results demonstrate that Spatial-SAE attains a 96% success rate on synthetic concept recovery tasks and significantly enhances the interpretability of DINOv2 activations.

0 citationsRead paper

When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task

Sep 29, 2026

This study investigates how large language models leverage representational geometry for computation in numerical comparison tasks. Through a causal geometric analysis of the Qwen model, this work reveals that numerical comparison is achieved via the coordinated interplay of attention mechanisms, residual connections, and MLP neurons, which collectively enable local interval comparison and maximum-value localization. The findings demonstrate that the model primarily relies on linear representations rather than manifold operations for decision-making, while theoretically establishing the coexistence of the manifold hypothesis and linear representational structures. By elucidating a multi-number decision mechanism grounded in linear superposition and local comparison, this research challenges the prevailing overreliance on curvature geometry within mechanistic interpretability, offering novel perspectives for understanding internal model computations.

0 citationsRead paper

Causal and Interpretable Structures in LLM Compositional Tasks

Sep 28, 2026

This study investigates how large language models represent and process relational information in compositional tasks. By analyzing the geometric structure and causal mechanisms of activations across Transformer layers, combined with prompt ensembling and causal intervention techniques, we conduct cross-model experiments on mainstream architectures such as Llama. Our findings reveal a hierarchical progression in recurrent conceptual reasoning, uncovering an incremental organizational mechanism wherein intermediate layers exploit binary relations while later layers leverage ternary relations, alongside the identification of causally inert structures. Furthermore, we demonstrate that constraining models to rely exclusively on causally relevant joint representations significantly improves next-token prediction accuracy.

0 citationsRead paper