REViT: Roto-reflection Equivariant Convolutional Vision Transformer

📅 2026-06-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing vision transformers struggle to effectively model rotational and reflectional symmetries in images, limiting their performance on orientation-sensitive tasks. This work proposes a discrete roto-reflection equivariant Vision Transformer that, for the first time, efficiently integrates discrete roto-reflection group equivariance into the ViT architecture. By combining convolutional attention mechanisms with equivariant convolutions, the method explicitly preserves rotational, reflectional, and positional symmetries during feature extraction. This approach substantially simplifies the implementation of equivariance while achieving superior performance over existing roto-reflection equivariant neural networks on image classification benchmarks.
📝 Abstract
In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs. In image classification and object detection, most of the studies on roto-reflection equivariant models have focused on using convolutional neural networks rather than vision transformers. In this paper, we examine the challenges involved in achieving equivariance in vision transformers, and we propose a simpler way to implement a discretized roto-reflection group equivariant vision transformer. The experimental results demonstrate that our approach outperforms the existing approaches for developing discrete roto-reflection group equivariant neural networks for image classification.
Problem

Research questions and friction points this paper is trying to address.

roto-reflection equivariance
vision transformer
convolutional attention
image classification
symmetry preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

roto-reflection equivariance
vision transformer
convolutional attention
group equivariant networks
discrete symmetry
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Sheir A. Zaheer
Sheir A. Zaheer
KAIST, S. Korea
Machine LearningAIRobotics
A
Alexander C. Holston
KC Machine Learning Lab, Seoul, Rep. of Korea
C
Chan Y. Park
KC Machine Learning Lab, Seoul, Rep. of Korea