🤖 AI Summary
This work addresses the limitations of Vision Transformers (ViTs) in dense prediction tasks, where low-resolution feature maps hinder performance and conventional upsampling methods often introduce artifacts such as feature leakage, fragmentation, and blurriness. The authors propose an implicit feature upsampling framework that operates without external image guidance, leveraging intermediate hidden states from the ViT to construct inter-layer continuous queries. This enables high-fidelity, task-agnostic feature prediction at arbitrary spatial coordinates while preserving feature-space alignment and supporting continuous-coordinate inference. Experimental results demonstrate significant improvements over existing image-guided upsampling approaches, with a +3.36 mIoU gain on Cityscapes semantic segmentation and a +8.09 increase in PCK@0.10 on the SPair-71k correspondence benchmark.
📝 Abstract
Vision Transformers (ViTs) have become a dominant architecture for visual representation learning, providing exceptionally strong and broadly reusable backbone features. However, ViTs are commonly operated on relatively small patch-token grids due to the quadratic cost of global self-attention, which creates a persistent bottleneck for dense prediction tasks such as semantic segmentation and depth estimation. This has motivated the development of task-agnostic feature upsamplers. While recent state-of-the-art methods produce visually sharp dense representations, their reliance on shallow image encoders for guided upsampling can introduce feature leakage, fragmentation, and blur. We introduce ViT-Up, an implicit feature upsampling framework that replaces external image guidance with layer-wise query construction from intermediate ViT hidden states. This enables feature prediction at arbitrary continuous image coordinates while preserving alignment with the backbone feature space. Experiments demonstrate that ViT-Up consistently outperforms state-of-the-art image-guided upsamplers across dense prediction and semantic correspondence. On DINOv3-S+, ViT-Up improves over prior methods by up to +2.07 mIoU on Cityscapes and +4.17 PCK@0.10 on SPair-71k. With the larger DINOv3-B backbone, these gains increase to +3.36 mIoU and +8.09 PCK@0.10, demonstrating that ViT-Up scales favorably with backbone capacity.