REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of achieving rotation- and reflection-equivariant feature extraction in large-scale Vision Transformers (ViTs). We propose a hierarchical equivariant ViT architecture based on windowed group convolutional self-attention. This work introduces the first scalable windowed group convolutional self-attention mechanism for large datasets, integrating rotation-reflection group equivariance theory with hierarchical network design to effectively overcome the parameter scalability bottleneck inherent in existing equivariant models. We successfully train and validate million-parameter equivariant models on ImageNet, demonstrating their practical viability at scale. Furthermore, all code and pretrained weights are publicly released. By providing an efficient and reproducible framework, this project establishes a new paradigm for large-scale equivariant visual representation learning.
📝 Abstract
We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.
Problem

Research questions and friction points this paper is trying to address.

equivariant vision transformer
roto-reflection equivariance
scalability
feature extraction
ImageNet
Innovation

Methods, ideas, or system contributions that make the work stand out.

Roto-reflection Equivariance
Group-Convolutional Self-Attention
Hierarchical Architecture
Vision Transformer
Scalability
🔎 Similar Papers
No similar papers found.