RLHarness: Co-evolving Procedural Skills with Reinforcement Learning for Long-horizon Multimodal Reasoning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges in multimodal long-chain reasoning where terminal verification struggles to identify reusable programs and static skill libraries misalign with learned policies. To overcome these limitations, this work proposes a unified versioned Harness framework that enables the co-evolution of skills and policies. The method alternately evolves external programs and reinforcement learning (RL) policies by integrating DAPO, supervised fine-tuning (SFT), exploratory distillation, and a post-RL reconstruction mechanism. This design facilitates the automatic repair and dynamic updating of programmatic skills, thereby transcending the constraints of fixed skill repositories. Experimental results demonstrate that the proposed paradigm achieves accuracies of 62% and 50% on the MetroMap and TravelMap benchmarks, respectively, alongside F1 scores of approximately 65% on Fee-VL and Cancel-VL, significantly enhancing long-chain reasoning performance.
📝 Abstract
Multimodal reasoning requires models to preserve visual evidence through long decision chains while selecting appropriate procedures across diverse scenarios and rules. When learning is guided only by terminal verifiers, reinforcement learning (RL) reveals whether a final answer is correct but not how it should be produced. The policy must therefore discover reusable reasoning procedures while learning to execute them, creating a program cold-start problem. Skills can externalize successful procedures, reduce repeated exploration, and provide inspectable guidance. However, a fixed Skill Bank assumes that this guidance remains compatible with an evolving policy, while updating Skills alone can leave their triggers, execution protocols, and demonstrations stale or mutually inconsistent. We introduce RLHARNESS, which organizes Skills, selection and execution protocols, few-shot demonstrations, and task contracts into a unified, versioned Harness and alternates Harness evolution with policy learning. An Exploration-Distillation Harness builds the initial Harness and version-aligned verified traces for SFT and DAPO I. After the first RL block, a Post-RL Reconstruction Harness rebuilds Skills, protocols, and demonstrations from fresh success-failure rollouts, and DAPO II adapts the policy to the reconstructed program. RLHARNESS improves Accuracy from 16.25%/27.50% to 62.00%/50.00% on MetroMap/TravelMap and raises F1 score from 37.13%/45.50% to 65.81%/65.51% on Fee-VL/Cancel-VL. All four tasks achieve their best results only after reconstruction and DAPO II, showing that an evolving Harness complements RL by continually updating the external program that the policy learns to execute.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Reasoning
Reinforcement Learning
Program Cold-start
Long-horizon Reasoning
Skill Co-evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Multimodal Reasoning
Co-evolution
Skill Bank
Program Synthesis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziqiao Shang
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China
Z
Zian Xu
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China
J
Ji-Chen Yan
Didichuxing Co. Ltd
W
Weiming Wu
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China
Z
Ziyi Jia
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China
J
Jie Meng
Didichuxing Co. Ltd
T
Tao Huang
Didichuxing Co. Ltd
S
Shan Huang
Didichuxing Co. Ltd
Lan-Zhe Guo
Lan-Zhe Guo
LAMDA Group, Nanjing University
Machine Learning