🤖 AI Summary
This study addresses the numerical instability and performance limitations of hyperbolic visual Transformers in ambient coordinates caused by large radii. To overcome these issues, this work proposes a fully hyperbolic Vision Transformer formulated in polar coordinates. Methodologically, it introduces a first-of-its-kind polar-coordinate fully connected layer that independently learns radius mappings, designs logarithmically growing hyperbolic displacements as relative positional encodings, and reformulates residual connections as Lorentz boosts based on the mean radius. These innovations effectively decouple directional and radial computations, mitigating numerical errors while fully exploiting the hierarchical properties of hyperbolic geometry. Experimental results demonstrate that the proposed method significantly outperforms both Euclidean and existing hyperbolic baselines on standard vision tasks, further exhibiting strong generalization capabilities on ImageNet.
📝 Abstract
Hyperbolic space can embed intrinsic hierarchies in data with low distortion due to the exponential growth of volume with distance from the origin. However, current Lorentz transformer blocks are formulated in ambient coordinates, where numerical errors increase at large radii due to instability in the Lorentzian inner product, leading many models to limit the radius to avoid this issue. This prevents us from utilizing the regions of hyperbolic space that motivate the geometry. To address this, we revisit the components of the transformer block in polar coordinates, where hyperbolic operations such as distance calculation and attention centroids can be computed without the numerical cancellation in their ambient-coordinate formulations. Specifically, we propose a polar fully connected layer that separately maps an embedding's direction and radius, allowing the radius of a feature to be learned rather than determined by the norm of a linear map. We additionally introduce horospherical shifts as relative positional encodings whose query-key distances grow logarithmically with the token gap. Finally, we reformulate the residual connection as average radius Lorentz boosts. Combining these components, we develop a fully hyperbolic transformer that substantially improves performance over Euclidean and hyperbolic baselines on standard vision tasks. We further evaluate our model on ImageNet and demonstrate its ability to generalize to other datasets using pre-trained weights, similar to Euclidean counterparts.