World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing vision-language-action (VLA) models, which fail to differentiate the roles of egocentric and wrist-mounted views in fine manipulation and lack explicit modeling of how wrist interactions evolve within task contexts. To bridge this gap, we propose W2-VLA, a novel framework that introduces, for the first time, a task-conditioned future wrist latent prediction mechanism to establish a compact interface between global task context and local wrist interaction. We further design a W2-CoT synthesis pipeline that generates structured supervision signals incorporating manipulation progress, physical transition cues, and wrist-based evidence, thereby enhancing semantic alignment of the latent variables. Experiments demonstrate that our approach significantly improves fine-grained and contact-sensitive manipulation performance for both single- and dual-arm systems on LIBERO, RoboTwin 2.0, and real-world tasks, achieving action generation rates exceeding 80 Hz.
📝 Abstract
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.
Problem

Research questions and friction points this paper is trying to address.

fine-grained manipulation
vision-language-action models
wrist-view prediction
task-conditioned modeling
robot manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-conditioned wrist modeling
vision-language-action (VLA)
future-aware manipulation
latent interface
fine-grained robot manipulation