Quantile Head for Vision-Language-Action Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of point regression, which discards distributional information in Vision-Language-Action (VLA) model action heads, and the high sampling costs associated with flow matching. To unify these objectives, we propose a quantile regression approach that employs a dedicated quantile head to predict ordered quantiles in a single forward pass, thereby recovering the complete action distribution. By adopting joint supervision to train a default median policy, the method remains compatible with diverse sampling strategies while significantly reducing update variance. Experimental evaluations on the LIBERO benchmark and real-world robotic tasks demonstrate that our approach achieves the highest average success rates and the shortest episode durations compared to existing baselines.
📝 Abstract
Vision-Language-Action (VLA) models integrate pretrained Vision-Language Models (VLMs) with action heads for robot control. Common action heads have distinct limitations: point regression provides only a point estimate of the action distribution, while standard flow-matching samplers require costly iterative sampling. To address these limitations, we unify regression and flow matching under a shared objective and extend it to derive a quantile objective. This quantile objective guides the design of our Quantile Head, which predicts a median and positive gaps to form ordered marginal action quantiles in one forward pass. These quantiles support multiple sampling strategies without retraining and are jointly supervised to train the default median policy. Our local analysis of this joint supervision shows that, with calibrated nearby quantiles, fixed gaps, and matched correction speed, direct median updates have lower variance than under median-only supervision. Experiments show that this jointly supervised median policy achieves the highest average success rates among the compared methods on LIBERO, LIBERO-Plus, LIBERO-Pro, and two real-robot tasks, together with the shortest mean episode time among matched LIBERO baselines; code is available at https://github.com/xwangrs/Quantile-Head-for-VLA.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
action head
point regression
flow matching
robot control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantile Head
Vision-Language-Action Models
Flow Matching
Joint Supervision
Robot Control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.