🤖 AI Summary
This study addresses the limitation imposed by fixed verifiers on policy self-improvement in embodied reasoning by proposing VeriFine, a framework that introduces a novel dual-loop co-evolutionary mechanism. Through the joint optimization of policy, curriculum, and verifier, VeriFine expands verification capabilities. By integrating adaptive curriculum generation with human-in-the-loop calibration, the framework leverages reinforcement learning, supervised fine-tuning, and rubric alignment techniques to dynamically overcome verification bottlenecks. Evaluated on driving and navigation tasks, VeriFine achieves sustained synergistic improvements in both policy performance and verification accuracy. This work establishes an effective paradigm for the self-evolution of embodied intelligence.
📝 Abstract
Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.