One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of single-block recurrent Transformers, which lack depth-specific transformation capabilities when reusing a shared block. To overcome this, we propose a feed-forward network (FFN) weight resampling mechanism based on a shared expert bank and normalized depth coordinates. This approach models the merged weight space as an optimal mixture-of-experts (MoE) family, supporting elastic-depth training and graph unfolding during deployment. By integrating knowledge distillation with weight-space interpolation, we construct a recurrent Vision Transformer that achieves full-depth accuracy parity with the DINOv2 teacher model while preserving its strong cross-task generalization. Notably, the proposed method reduces the parameter count by 70%, offering a highly efficient yet performant architecture for vision tasks.
📝 Abstract
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
Problem

Research questions and friction points this paper is trying to address.

Recurrent Vision Transformer
Parameter Efficiency
Depth-Programmed Experts
Elastic Depth
Mixture of Experts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recurrent Vision Transformers
Mixture of Experts
Weight-space Merging
Elastic-depth Training
Knowledge Distillation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.