Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current world action models suffer from entangled representations of high-level physical state evolution and low-level action trajectories, which constrains both world modeling capacity and action generation performance. This work proposes PILOT, a novel framework that achieves the first explicit disentanglement of intent and trajectory within world action models. PILOT introduces a Representation Deduction (RD) mechanism to explicitly model latent state-transition tokens and incorporates a motion chain-of-thought (CoT) as an intrinsic capability to guide fine-grained trajectory generation. This design enhances physical interpretability and leverages state-transition supervision to alleviate the sparse reward problem in action generation. Experiments demonstrate that PILOT significantly improves success rates and generalization on complex robotic manipulation tasks, enables efficient few-shot fine-tuning on real robots, and integrates seamlessly into mainstream world action model architectures.
📝 Abstract
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical condition evolution and low-level action trajectory generation within the Action Model, creating a structural bottleneck while weakening the predictive capability of world evolution modeling for action generation. We propose PILOT (Physical Inference for Latent Optimized Trajectories), whose core Representational Deduction (RD) bridges this gap by integrating motion thought-of-chain (CoT) guidance as a native model capability. Specifically, RD aims to encourage the action branch to explicitly model potential state transition tokens, which are retained as CoT in the reasoning space to guide fine-grained motion trajectory. Experiments demonstrate that RD not only significantly improves the success rate and generalization ability of WAMs in complex robotic manipulation tasks but also enhances the model's physical interpretability by decoupling high-level motion semantics from low-level trajectory details. Furthermore, the abundant state transition supervision signals introduced by RD effectively alleviate the sparse supervision in action generation, enabling it to serve as an efficient few-shot real-robot fine-tuning strategy and demonstrating superior scalability for migration to mainstream WAM architectures.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
representational entanglement
state transition
motion trajectory
action generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Representational Deduction
World Action Models
Chain-of-Thought
State Transition Modeling
Trajectory Decoupling
🔎 Similar Papers
No similar papers found.