Cue the Flow: Steering Flow-Matching Policies for Open-World Delivery Manipulation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient robustness of low-level controllers in open-world delivery under noisy perception and dynamic environments by proposing a bottom-level control framework built upon pretrained flow-matching vision-language-action (VLA) models. The method introduces a lightweight cue-conditioned adapter that translates visual prompts into spatial guidance signals, alongside a contrastive learning-based cue encoding mechanism coupled with diagonal affine transformation for action alignment, enabling efficient integration of VLA priors with target information. Experimental results demonstrate that the proposed approach substantially enhances instruction-following capabilities across both tabletop and mobile-base scenarios, nearly doubling the average task success rate while maintaining minimal inference overhead.
📝 Abstract
Open-world goods delivery requires mobile manipulators to follow free-form user instructions and manipulate potentially novel objects. Existing dual-system approaches use high-level grounding models to convert language into grounded visual prompts, but their low-level controllers can remain brittle under noisy perception, dynamic scenes, and contact-rich interactions. We instead use a pretrained flow-matching vision-language-action model as the low-level control interface, leveraging its reactivity and robustness to environmental changes while treating the grounding output as a spatial cue for policy steering. Our key insight is that the pretrained VLA already provides a strong manipulation prior, while the spatial cue supplies the missing target information needed to guide actions under novel language--object mappings. Concretely, we introduce a lightweight cue-conditioned adapter. The adapter is first trained with contrastive objectives to produce salient and spatially discriminative cue representations, and is then supervised to predict a diagonal affine transformation over the generated action chunk, aligning policy steering with the cued target. Across tabletop and mobile-base settings, our method improves instruction following and manipulation success on both in-domain and out-of-domain objects, achieving up to near $2\times$ improvement in average task success rate with negligible inference overhead.
Problem

Research questions and friction points this paper is trying to address.

open-world manipulation
vision-language-action model
mobile manipulation
instruction following
policy steering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-Matching VLA
Cue-Conditioned Adapter
Policy Steering
Contrastive Learning
Open-World Manipulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Haoxuan Wang
Haoxuan Wang
PhD, University of Illinois Chicago
Machine Learning Efficiency
G
Griffin Galimi
University of California, Los Angeles
J
Junhua Huang
University of California, Los Angeles
S
Selina Song
University of California, Los Angeles
W
Wayne Wu
University of California, Los Angeles
Yan Yan
Yan Yan
University of Illinois Chicago
Computer VisionMultimediaMachine Learning
Bolei Zhou
Bolei Zhou
Associate Professor at UCLA
Computer VisionRoboticsArtificial Intelligence