VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses signal collapse and difficulty mismatch arising from static environments in reinforcement learning for vision-language models by proposing VICO, a co-evolutionary framework. VICO introduces a novel environment-agent co-evolution mechanism grounded in structured image editing—such as scene graphs and charts—rather than textual modification. By employing an environment rewriter to dynamically generate adaptive training samples, combined with pass-rate-driven difficulty calibration and reinforcement learning with verifiable rewards (RLVR), the framework achieves self-adaptive training without requiring additional annotations. Evaluated across nine multimodal benchmarks, VICO yields an average out-of-domain performance improvement of 5.0%, surpassing baseline methods by 4.3% to 8.4%. Notably, it attains performance comparable to specialized approaches while utilizing minimal annotated data.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Reinforcement Learning
Post-training
Visual Reasoning
Static Training Environment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Co-evolutionary framework
Reinforcement learning with verifiable rewards (RLVR)
Environment-as-Rewriter
Vision-language models
Visual reasoning