PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge in reinforcement fine-tuning of Vision-Language-Action (VLA) models, where sparse rewards hinder long-horizon task completion. To this end, we propose Progress Field Reinforcement Learning, which models task progress as a geometric structure. By replacing scalar predictions with geometric distances to construct goal-conditioned value spaces, our method generates dense feedback signals for fine-grained optimization. Technically, it leverages pretrained VLA features alongside a shared progress field head, trained jointly through offline and online reinforcement learning under geometric constraints. Extensive experiments on the LIBERO benchmark and real-world bimanual manipulation tasks demonstrate that the proposed approach significantly outperforms existing supervised and reinforcement learning baselines.
📝 Abstract
Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progress and use it as dense feedback for policy improvement. Despite their architectural differences, existing progress-aware methods commonly formulate task progress as an explicit scalar prediction, providing limited structure for modeling how intermediate observations relate to the task goal, which may hinder effective transition-level credit assignment. We introduce Progress Field Reinforcement Learning (PF-RL), which learns a structured goal-conditioned progress representation over pretrained VLA features and converts it into dense credit for policy optimization. A lightweight shared Progress Field head maps current and goal representations into a compact progress space, where geometric distance induces goal-conditioned value, while complementary temporal and goal-structure objectives shape the learned geometry. Transition-level value changes naturally yield dense progress advantages, enabling fine-grained credit assignment for both offline policy improvement and online reinforcement fine-tuning. Extensive experiments on LIBERO, RoboTwin2.0, and real-world bimanual manipulation tasks show that PF-RL consistently improves policy performance over strong supervised fine-tuning, reinforcement fine-tuning, and progress-aware baselines.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Fine-Tuning
Vision-Language-Action Models
Sparse Reward
Credit Assignment
Task Progress Modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Progress Field Reinforcement Learning
Goal-Conditioned Value Geometry
Vision-Language-Action Models
Dense Credit Assignment
Reinforcement Fine-Tuning
🔎 Similar Papers
No similar papers found.