TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the degradation of reasoning capabilities in vision-language models (VLMs) during test-time reinforcement learning, where shared perceptual errors and sequence-level rewards fail to localize visual bottlenecks. To overcome this, we propose a multi-perspective self-distillation framework with visual contrastive token selection that decouples update directions from spatial positions. By integrating answer-level self-distillation, contrastive token filtering, and group relative advantage policy gradient, our method precisely updates visually sensitive regions without requiring labeled data, external verifiers, or teacher models. Extensive experiments across three VLMs and seven benchmarks demonstrate significant performance improvements, notably achieving a 13.53% accuracy gain on MMMMU for InternVL3-2B. This work enables efficient adaptation with minimal samples while preserving reasoning integrity.
πŸ“ Abstract
Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.
Problem

Research questions and friction points this paper is trying to address.

Test-time reinforcement learning
Vision-language models
Perceptual errors
Sequence-level rewards
Unlabeled adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Reinforcement Learning
Self-Distillation
Vision-Language Models
Visual Contrastive Token Selection
Policy Gradient
πŸ”Ž Similar Papers
2024-03-04Computer Vision and Pattern RecognitionCitations: 3