ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of geometric information caused by text-centric representations in existing skill agents and the disconnect between skill construction and policy optimization. To this end, we propose a vision-native skill learning framework that pioneers encoding successful interactions into visual skill cards. By integrating vision-language models with reinforcement learning, the framework enables retrieval-guided reasoning and reward shaping, while introducing a cold-start mechanism to accelerate early exploration. This design establishes a closed-loop feedback system for simultaneous skill accumulation and policy optimization. Experimental results demonstrate that the proposed method achieves a success rate of 0.91 across multiple benchmarks, comprehensively outperforming existing baselines and exhibiting significantly faster convergence than standard Proximal Policy Optimization algorithms.
📝 Abstract
Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming a closed feedback loop in which skill accumulation and policy improvement reinforce each other. An optional cold-start mechanism further accelerates early-stage learning. Evaluated on Sokoban, FrozenLake, and PrimitiveSkill, ViSkill achieves an overall success rate of 0.89, rising to 0.91 with cold-start initialization, outperforming all evaluated proprietary and open-source baselines while converging faster than standard PPO. Our code is available at https://github.com/ZJU-REAL/ViSkill.
Problem

Research questions and friction points this paper is trying to address.

skill-augmented agents
visual-native skills
VLM agents
policy optimization
sample efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual-native skills
VLM agents
Reward shaping
Closed feedback loop
Cold-start mechanism
💼 Related Jobs
No related jobs found.
H
Hongxing Li
Zhejiang University
D
Dingming Li
Zhejiang University
Yixin Li
Yixin Li
Stony Brook University
PET InstrumentMedical ImagingX-ray Imaging
Y
Yong Du
Zhejiang University
Wenqi Zhang
Wenqi Zhang
Zhejiang University
Language ModelMultimodal LearningEmbodied Agents
Weiming Lu
Weiming Lu
Zhejiang University
Natural Language ProcessingLarge Language ModelsAGI
J
Jun Xiao
Zhejiang University
Y
Yueting Zhuang
Zhejiang University
Y
Yongliang Shen
Zhejiang University