Symmetry-Aware Feature Learning: A Polynomial Separation for Multi-Index Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear sample complexity gap between symmetry-aware and symmetry-agnostic feature learning in multi-index models. By employing spherical online stochastic gradient descent, specific correlation loss functions, and convolutional network architectures, it systematically compares the directional recovery efficiency of weight sharing, data augmentation, and symmetry-agnostic learning under high-dimensional covariates. The primary contribution is establishing the first theoretical bounds demonstrating a polynomial separation in sample complexity between these two learning paradigms. Both theoretical analysis and empirical results reveal that full-group data augmentation achieves sample efficiency comparable to weight sharing, with both approaches significantly outperforming symmetry-agnostic training methods.
📝 Abstract
We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in $\mathbb{R}^d$ and $r=Θ(d^δ)$ teacher directions forming a cyclic symmetry orbit, where $0<δ<1/2$. We compare three ways of exploiting this structure: architectural weight sharing, data augmentation over the full symmetry group, and learning without access to the symmetry. In particular, we analyze a symmetry-tied convolutional network, an untied network, and the same untied network trained with full-group data augmentation, using spherical online SGD with correlation loss. For a class of polynomial links with information exponent $p\ge3$, we prove matching sample complexity bounds up to logarithmic factors: the tied and augmented learners achieve weak directional recovery in $\widetildeΘ(d^{p-1})$ samples, whereas the symmetry-agnostic learner requires $\widetildeΘ(rd^{p-1})$. For the pure quadratic Hermite link, the same separation holds for weak recovery of the teacher subspace, with sample complexities $\widetildeΘ(d)$ and $\widetildeΘ(rd)$, respectively. Thus, full-group data augmentation matches the sample efficiency of architectural weight sharing, and both provide a polynomial advantage over training without symmetry. For $p\ge3$, the proof reveals a two-stage mechanism: fluctuations at initialization select one direction in the teacher orbit, after which localized growth amplifies its overlap to the weak recovery scale while competing overlaps remain near their initialization scale.
Problem

Research questions and friction points this paper is trying to address.

multi-index models
symmetry-aware feature learning
sample complexity separation
data augmentation
weight sharing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Symmetry-Aware Feature Learning
Multi-Index Models
Sample Complexity Separation
Data Augmentation
Weight Sharing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jivan Waber
Information, Learning and Physics Laboratory, École Polytechnique Fédérale de Lausanne (EPFL)
V
Vanessa Piccolo
Information, Learning and Physics Laboratory, École Polytechnique Fédérale de Lausanne (EPFL)
Yatin Dandi
Yatin Dandi
EPFL, IIT Kanpur
Deep Learning TheoryStatistical PhysicsOptimization
Florent Krzakala
Florent Krzakala
École polytechnique fédérale de Lausanne
Statistical MechanicsStatisticsMachine LearningInformation theorySpin Glasses