NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of cross-platform data scaling and insufficient physical consistency in language-conditioned robotic manipulation by proposing the NarrativeFlow framework. This method models robot trajectories as embodiment-agnostic continuous velocity fields, overcoming the limitations of sparse keypoint approximations. Building upon vision-language-action models, it integrates flow matching with language-conditioned generation to synthesize physically plausible manipulation actions. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines on standard benchmarks and substantially improves success rates across diverse real-world tasks. These findings validate the effectiveness of NarrativeFlow in facilitating efficient cross-platform data reuse for language-conditioned robotic manipulation.
📝 Abstract
We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains limited by the labor-intensive collection of embodiment-specific data. Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language. Accordingly, NarrativeFlow generates robot flows that are physically consistent with real-world manipulation. To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks. The project page is available at https://shota0520.github.io/NarrativeFlow-project-page/
Problem

Research questions and friction points this paper is trying to address.

language-conditioned manipulation
robot velocity fields
embodiment-agnostic representation
robot foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-Matching
Velocity Fields
Vision-Language-Action Model
Embodiment-Agnostic
Language-Conditioned Manipulation
💼 Related Jobs
No related jobs found.
S
Shota Kobayashi
Keio University, Japan
K
Koki Seno
Keio University, Japan
D
Daichi Yashima
Keio University, Japan
Komei Sugiura
Komei Sugiura
Professor, Keio University
Multimodal AIRobot LearningEmbodied AIMachine Learning