PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vanishing gradient problem and coarse-grained credit assignment in reinforcement learning for multi-turn vision-language agents by proposing the PIVOT framework. PIVOT internalizes pivot localization and state recovery into parameter updates, thereby eliminating dependence on environment rollback. It is the first to reveal the physical rollback mechanism, enabling efficient self-distillation through a unified architecture. Furthermore, the framework integrates techniques including GRPO, confidence-gated OPD, visual trajectory tiling, and counterfactual rollback probes. Experimental results demonstrate that PIVOT achieves an accuracy of 0.90 across five benchmarks, yielding an 8% improvement over baselines and surpassing existing state-of-the-art methods by 5%.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) via Group-Relative Policy Optimization (GRPO) is widely used for multi-turn VLM agent training, yet it suffers from zero-gradient silence on uniform failures and coarse episode-level credit assignment. While On-Policy Distillation (OPD) and On-Policy Self-Distillation (OPSD) mitigate sparse rewards using hindsight information, their underlying mechanisms remain poorly understood. Through controlled counterfactual rollback probes across five multi-turn VLM agent benchmarks, we reveal that performance gains in OPSD/OPD are largely driven by physical state rollback at the pivot step, defined as the first unrecoverable action without remaining step budget. However, physical state rollbacks are computationally prohibitive and infeasible in real-world environments. To bridge this gap, we present Pivot-Aware Internalized Visual On-Policy Training (PIVOT), an RL framework that internalizes pivot localization and state restoration directly into token-level parameter updates, eliminating environment rollbacks during RL training and additional skill hints at test time. PIVOT unifies three functional roles within a single architecture: a failure Analyzer non-invasively localizes the pivot step and diagnoses failure modes from visual trajectory collages and action logs; a detached Teacher re-scores failed tokens under this privileged diagnostic context; and a Student optimizes joint GRPO and confidence-gated OPD objectives. At test time, both Teacher and Analyzer branches are stripped. Evaluated on five multi-turn VLM agent tasks across cognitive grid puzzles, 3D embodied control and navigation, and generative reasoning, PIVOT achieves 0.90 overall accuracy on Qwen2.5-VL-3B (+8% over SFT+GRPO baseline and +5% over previous SOTA) and scales to 0.92 on Qwen3-VL-2B (+12% over SFT+GRPO baseline).
Problem

Research questions and friction points this paper is trying to address.

multi-turn VLM agents
reinforcement learning
credit assignment
zero-gradient silence
state rollback
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Pivot-Aware
Multi-Turn VLM Agents
GRPO
Token-Level Optimization
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1