🤖 AI Summary
This study investigates how large language models leverage representational geometry for computation in numerical comparison tasks. Through a causal geometric analysis of the Qwen model, this work reveals that numerical comparison is achieved via the coordinated interplay of attention mechanisms, residual connections, and MLP neurons, which collectively enable local interval comparison and maximum-value localization. The findings demonstrate that the model primarily relies on linear representations rather than manifold operations for decision-making, while theoretically establishing the coexistence of the manifold hypothesis and linear representational structures. By elucidating a multi-number decision mechanism grounded in linear superposition and local comparison, this research challenges the prevailing overreliance on curvature geometry within mechanistic interpretability, offering novel perspectives for understanding internal model computations.
📝 Abstract
One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.