🤖 AI Summary
This study addresses the polar distortion and coordinate singularities in existing spherical Transformers caused by the neglect of underlying geometric structures. To overcome these limitations, this work proposes a relative positional encoding scheme based on unitary SO(3) representations that rigorously incorporates spherical geometry into the attention mechanism, yielding an architecture that is both equivariant and compatible with FlashAttention. Evaluated on the rotating sphere shallow water dynamics prediction task, the proposed method achieves lower prediction errors and reduced inference times compared to the S2Transformer baseline. These results demonstrate that the approach effectively reconciles geometric rigor in spherical modeling with computational efficiency.
📝 Abstract
Spherical data arise in many scientific applications. Often spherical transformers disregard the geometry of the underlying spherical domain, causing distortions and coordinate singularities near the poles. We introduce SO(3)-RoPE, a relative positional embedding that incorporates spherical geometry into transformer attention through unitary SO(3) representations. Our formulation is SO(3)-equivariant and compatible with FlashAttention, retaining efficiency of vanilla transformers. On shallow water dynamics prediction over a rotating sphere, our SO3ViT outperforms an S2Transformer baseline with lower errors and reduced runtime.