How Do Transformers Learn to Represent Symmetries?

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the mechanisms by which Transformers learn geometric symmetries of point clouds under limited data augmentation. Through systematic analysis of internal model structures, it reveals their capacity to learn and extrapolate across different symmetry groups. Notably, this work is the first to identify a hierarchical order in symmetry learning, uncovering a pattern of increasing learnability progressing from non-angle-preserving to fundamental angle-preserving subgroups. Furthermore, it elucidates that the interpretable mechanism inducing approximate invariance serves as a critical building block for equivariant learning. Overall, this research provides new perspectives on understanding the geometric inductive biases of Transformers and effectively advances the learning of equivariant functions.
📝 Abstract
Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the Transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we study the ability of a vanilla Transformer to learn various symmetries through finite data augmentation for point cloud datasets. We identify an ordering of increasing learnability across the following symmetry groups: (i) non-angle-preserving symmetries, (ii) angle-preserving symmetries, and (iii) base angle-preserving subgroups, such as translation, rotation, and scale. For the base angle-preserving groups, we further investigate the Transformer's extrapolation behavior and conduct a structural analysis of the trained models, allowing us to identify interpretable mechanisms that induce invariance. Finally, we extend our analysis to equivariant functions and show that the detected mechanisms for approximate invariance can also provide a key building block for learned equivariance. Our project page is available at https://transformers-learn-symmetries.github.io/
Problem

Research questions and friction points this paper is trying to address.

Transformers
symmetry learning
data augmentation
point clouds
invariance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer
symmetry learning
data augmentation
point cloud
equivariance
🔎 Similar Papers
E
Eduardo Santos-Escriche
Technical University of Munich, Munich Center for Machine Learning
V
Valerie Engelmayer
Technical University of Munich, Munich Center for Machine Learning
Ya-Wei Eileen Lin
Ya-Wei Eileen Lin
PhD at the Technion
Geometric LearningGraph Machine Learning
Stefanie Jegelka
Stefanie Jegelka
TUM and MIT
Machine LearningOptimizationSubmodularityGraph Neural Networks