SO(3)-RoPE for Spherical Transformers
This study addresses the polar distortion and coordinate singularities in existing spherical Transformers caused by the neglect of underlying geometric structures. To overcome these limitations, this work proposes a relative positional encoding scheme based on unitary SO(3) representations that rigorously incorporates spherical geometry into the attention mechanism, yielding an architecture that is both equivariant and compatible with FlashAttention. Evaluated on the rotating sphere shallow water dynamics prediction task, the proposed method achieves lower prediction errors and reduced inference times compared to the S2Transformer baseline. These results demonstrate that the approach effectively reconciles geometric rigor in spherical modeling with computational efficiency.