🤖 AI Summary
This study addresses the high computational overhead of self-attention in Vision Transformers (ViTs) and the lack of online adaptability in existing pruning methods by proposing the DORA framework. This approach formulates token pruning as a finite-horizon Markov decision process, leveraging deep reinforcement learning to generate input-adaptive, layer-wise dynamic pruning policies for frozen ViTs. To facilitate efficient training, DORA introduces a hierarchical actor, a privileged critic, and a shadow evaluation credit assignment mechanism. Evaluated on ImageNet, DORA reduces FLOPs by 38.4% with less than 1% accuracy degradation while improving throughput by 32.4%. Furthermore, the proposed method demonstrates significant zero-shot transfer capabilities, highlighting its effectiveness and generalizability across diverse downstream tasks without requiring task-specific fine-tuning.
📝 Abstract
Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose DORA (Dynamic Online Reinforcement Agent), which learns an input-adaptive pruning policy itself for frozen ViTs. At each eligible block, a hierarchical actor decides whether to prune, how many tokens to remove, and which tokens to remove from each image's evolving representation. Because early deletions change the states observed by later decisions, DORA formulates pruning as a finite-horizon Markov decision process. Complete-prefix shadow evaluations convert final-prediction fidelity into localized per-step credit, while closed-loop accuracy feedback adjusts the fidelity penalty toward a shared accuracy-drop target. A privileged critic and all shadow computations are training-only. Deployment retains the frozen backbone and a lightweight actor that applies hard deletion and packed variable-length FlashAttention, converting token reduction into measured speedups. On ImageNet-1K with DeiT-Base, DORA reduces FLOPs by 38.4% relative to the uncompressed backbone within one percentage point of accuracy loss. Averaged across four ViT-type backbones at matched accuracy, DORA uses 13.2% fewer FLOPs and achieves 32.4% higher throughput than the corresponding per-backbone baseline means. Under zero-shot transfer to ImageNet-A, these gains widen to 20.3% and 45.6%, respectively.