SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that multimodal agent training suffers from sparse rewards, leaving intermediate decisions unsupervised, while manual evaluation criteria are difficult to scale and disconnected from execution. To overcome these limitations, this work proposes SkillRubric, a framework that leverages structured skills to unify action guidance and success criteria. It introduces an alternating co-evolution mechanism that revises guidance under frozen policies and optimizes evaluation criteria offline. Furthermore, a vision-language model-based multimodal verifier is employed to assign fine-grained rewards by integrating screenshots with tool outputs. The primary contribution lies in achieving the co-evolution of guidance and evaluation, yielding consistent performance improvements across multiple multimodal agent benchmarks. These results demonstrate that evolved skills provide superior planning and tool-use instructions for agent training.
📝 Abstract
Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervision for intermediate decisions. Rubric-based rewards address this limitation through explicit intermediate criteria, but reliable rubrics are difficult to construct at scale and often disconnected from the procedure followed by the actor. We observe that a well-structured skill naturally specifies both how to act and what successful execution should achieve. Based on this insight, we introduce SkillRubric, which represents each skill through aligned actor-facing guidance and an evaluator-facing rubric. A multimodal verifier evaluates skill-defined goals using screenshots and tool outputs, assigning completion and progress rewards to the responsible turns. We further introduce an alternating co-evolution scheme that validates guidance revisions through paired rollouts under a frozen policy and rubric revisions offline under fixed guidance. Experiments across diverse multimodal agent benchmarks demonstrate consistent performance gains, while controlled paired rollouts further show that evolved skills provide more effective guidance for planning and tool use than their preceding versions.
Problem

Research questions and friction points this paper is trying to address.

multimodal agents
sparse rewards
rubric-based evaluation
skill guidance
intermediate supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Agents
Skill Co-evolution
Rubric-based Rewards
Process Supervision
Alternating Optimization