VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underutilization of complementary visual information in multi-agent reinforcement learning by proposing the CVED and HAPI frameworks. Specifically, CVED employs on-policy distillation to transform parallel trajectory observations into shared supervisory signals, thereby facilitating the internalization of visual experience and aligning observation with decision-making. Complementarily, HAPI optimizes heterogeneous perception policies through a combination of success reinforcement and failure guidance mechanisms. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across both fine-grained perception and general reasoning tasks, significantly outperforming existing baseline models.
📝 Abstract
Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.
Problem

Research questions and friction points this paper is trying to address.

active multimodal agents
reinforcement learning
collective visual experience
on-policy distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Collective Visual Experience
Active Multimodal Agents
Heterogeneity-Aware Policy Improvement
Experience-Conditioned Teacher
🔎 Similar Papers
No similar papers found.
Z
Zheng Jiang
Tsinghua University
H
Houde Qian
Tsinghua University
Yiming Chen
Yiming Chen
Tsinghua University
deep learning
L
Ling Li
Tsinghua University
C
Chaoyang Li
Tsinghua University
Y
Yueqi Li
Tsinghua University
Y
Yuxuan Liu
Tsinghua University
L
Lifeng Sun
Tsinghua University