Objective-Specific Privileged Bases via Full-Prefix Matryoshka Learning

📅 2026-05-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing representation learning methods suffer from non-identifiable dimensions and a lack of semantic ordering under rotational invariance, making it difficult to align representations with downstream task objectives. This work proposes Full-Prefix Matryoshka Representation Learning (MRL) and, for the first time in a linear setting, theoretically demonstrates its ability to recover principal directions ordered by task relevance, thereby constructing a task-aligned semantic basis. Theoretical analysis reveals that the magnitudes of embedding coordinates effectively reflect the informational importance of individual dimensions. Coupled with an efficient computation mechanism leveraging shared statistics, the learned representations exhibit stable, task-oriented structural properties in empirical evaluations.
📝 Abstract
Learned representations are often invariant to rotational transformations, leaving individual dimensions non-identifiable and interchangeable. We study how Matryoshka Representation Learning (MRL) induces a task-aligned privileged basis distinct from variance-based or regularizer-induced orderings. In the linear setting, we prove that full-prefix MRL recovers the ordered principal directions, and can be computed efficiently using shared statistics. Empirically, we demonstrate that MRL yields consistent per-dimension structure aligned with task signal, where coordinate magnitude reflects informativeness.
Problem

Research questions and friction points this paper is trying to address.

Matryoshka Representation Learning
privileged basis
rotational invariance
task-aligned representation
dimension identifiability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Matryoshka Representation Learning
privileged basis
full-prefix learning
task-aligned representation
dimensional identifiability