Institution profile

Laboratoire de Mécanique des Solides

Academic institutioneurope · fr
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Predictive Geometry of Hidden Trajectories in Transformers

Sep 29, 2026

This study addresses the unclear geometric properties of hidden states in Transformers trained solely with terminal losses, particularly regarding their constraints from downstream computation. We analyze the local second-order geometry of layer-wise losses and employ the pullback Fisher operator to identify output-sensitive directions and predictive null spaces, thereby constructing an observable residual stream subspace. Furthermore, we propose a token-level curvature score based on Fisher-weighted sensitivity, serving as a loss-aware alternative to attention magnitude for enabling non-uniform hierarchical rank allocation, which is efficiently estimated via matrix-free Jacobian-vector products. Evaluated on datasets such as WikiText, this approach effectively predicts perturbation sensitivity, provides competitive structured token pruning signals, and significantly enhances the recovery performance of low-rank student models during autoregressive distillation.

0 citationsRead paper

The Geometry of Inference in Transformer Residual Streams

Sep 29, 2026

This study investigates how representations within the residual stream of Transformers evolve across layers to yield specific predictions. To this end, it constructs a high-dimensional geometric framework that contrasts cosine and Euclidean distances, empirically analyzing the norms, alignments, and endpoint geometries of intermediate and final states. The authors rigorously prove that linear trajectories introduce no new competitors and reveal how directional shifts precipitate a sharp reduction in competition. Furthermore, this work quantifies the growth of geometric specificity during inference, establishing a theoretical link between the geometric structure of the residual stream and the organization of model outputs.

0 citationsRead paper

Retrieval Capacity of Self-Attention Under Competition

Sep 29, 2026

This study addresses the unclear extent to which language models effectively utilize contextual tokens through self-attention and the underlying determinants thereof. We propose a retraining-free metric for effective attention set size, estimated by retaining high-attention-weight tokens and measuring changes in negative log-likelihood. Combined with conditional theoretical modeling, this approach analyzes how effective sets evolve with context length and competition dynamics. Our findings reveal how attention competition and aggregation influence retrieval capacity. Furthermore, we demonstrate that small effective attention sets suffice to maintain low loss, that attention-based token selection significantly outperforms random baselines, and that weight normalization substantially reduces the required set size.

0 citationsRead paper
Recent publications

Latest Papers

Predictive Geometry of Hidden Trajectories in Transformers

Sep 29, 2026

This study addresses the unclear geometric properties of hidden states in Transformers trained solely with terminal losses, particularly regarding their constraints from downstream computation. We analyze the local second-order geometry of layer-wise losses and employ the pullback Fisher operator to identify output-sensitive directions and predictive null spaces, thereby constructing an observable residual stream subspace. Furthermore, we propose a token-level curvature score based on Fisher-weighted sensitivity, serving as a loss-aware alternative to attention magnitude for enabling non-uniform hierarchical rank allocation, which is efficiently estimated via matrix-free Jacobian-vector products. Evaluated on datasets such as WikiText, this approach effectively predicts perturbation sensitivity, provides competitive structured token pruning signals, and significantly enhances the recovery performance of low-rank student models during autoregressive distillation.

0 citationsRead paper

The Geometry of Inference in Transformer Residual Streams

Sep 29, 2026

This study investigates how representations within the residual stream of Transformers evolve across layers to yield specific predictions. To this end, it constructs a high-dimensional geometric framework that contrasts cosine and Euclidean distances, empirically analyzing the norms, alignments, and endpoint geometries of intermediate and final states. The authors rigorously prove that linear trajectories introduce no new competitors and reveal how directional shifts precipitate a sharp reduction in competition. Furthermore, this work quantifies the growth of geometric specificity during inference, establishing a theoretical link between the geometric structure of the residual stream and the organization of model outputs.

0 citationsRead paper

Retrieval Capacity of Self-Attention Under Competition

Sep 29, 2026

This study addresses the unclear extent to which language models effectively utilize contextual tokens through self-attention and the underlying determinants thereof. We propose a retraining-free metric for effective attention set size, estimated by retaining high-attention-weight tokens and measuring changes in negative log-likelihood. Combined with conditional theoretical modeling, this approach analyzes how effective sets evolve with context length and competition dynamics. Our findings reveal how attention competition and aggregation influence retrieval capacity. Furthermore, we demonstrate that small effective attention sets suffice to maintain low loss, that attention-based token selection significantly outperforms random baselines, and that weight normalization substantially reduces the required set size.

0 citationsRead paper