Institution profile

abaka

Industry research
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

Oct 07, 2026

This study addresses signal collapse and difficulty mismatch arising from static environments in reinforcement learning for vision-language models by proposing VICO, a co-evolutionary framework. VICO introduces a novel environment-agent co-evolution mechanism grounded in structured image editing—such as scene graphs and charts—rather than textual modification. By employing an environment rewriter to dynamically generate adaptive training samples, combined with pass-rate-driven difficulty calibration and reinforcement learning with verifiable rewards (RLVR), the framework achieves self-adaptive training without requiring additional annotations. Evaluated across nine multimodal benchmarks, VICO yields an average out-of-domain performance improvement of 5.0%, surpassing baseline methods by 4.3% to 8.4%. Notably, it attains performance comparable to specialized approaches while utilizing minimal annotated data.

0 citationsRead paper
Recent publications

Latest Papers

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

Oct 07, 2026

This study addresses signal collapse and difficulty mismatch arising from static environments in reinforcement learning for vision-language models by proposing VICO, a co-evolutionary framework. VICO introduces a novel environment-agent co-evolution mechanism grounded in structured image editing—such as scene graphs and charts—rather than textual modification. By employing an environment rewriter to dynamically generate adaptive training samples, combined with pass-rate-driven difficulty calibration and reinforcement learning with verifiable rewards (RLVR), the framework achieves self-adaptive training without requiring additional annotations. Evaluated across nine multimodal benchmarks, VICO yields an average out-of-domain performance improvement of 5.0%, surpassing baseline methods by 4.3% to 8.4%. Notably, it attains performance comparable to specialized approaches while utilizing minimal annotated data.

0 citationsRead paper