Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear semantic division of labor between register tokens and high-norm outlier tokens in Vision Transformers. To investigate this, we employ sparse autoencoders integrated with an automated interpretability pipeline, UMAP clustering, and causal ablation experiments to quantitatively analyze internal token representations within DINOv2. Our work provides the first empirical evidence of a significant functional asymmetry between these two token types: register tokens encode high-level semantics, whereas outlier tokens predominantly capture low-level textures. Notably, ablating register features results in a 48.17% precipitous decline in representation similarity. These findings offer critical empirical support for understanding token specialization mechanisms in self-supervised Vision Transformers.
📝 Abstract
Self-supervised Vision Transformers (ViTs), such as DINOv2, learn rich visual representations, but the functions of their internal tokens remain poorly understood. Recent architectures introduce dedicated register tokens to reduce high-norm out- lier patch tokens that emerge in background re- gions, yet the semantic and functional roles of both token types have not been fully established. In this paper, we analyze these roles by training sparse autoencoders (SAEs) on register-token and outlier-token activations in DINOv2. Using an automated interpretability pipeline, UMAP clus- tering, and CLIP-space cross-checks, we find that register-token features are more strongly associ- ated with high-level semantic concepts. Outlier- token features, by contrast, are more often associ- ated with lower-level structural, background, and texture-dominant patterns. Causal ablations fur- ther reveal a substantial functional asymmetry: disrupting top-activating register-derived features produces a 48.17% drop in representation cosine similarity, whereas disrupting outlier-derived fea- tures produces only a 0.31% drop. Together, our results provide evidence for token specialization in self-supervised ViTs.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
register tokens
outlier patch tokens
self-supervised learning
token interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Transformers
Sparse Autoencoders
Register Tokens
Mechanistic Interpretability
Self-supervised Learning
🔎 Similar Papers
No similar papers found.