Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of explicit long-horizon intent in generative robot policies and the prohibitive cost of existing solutions by proposing a proprioceptive action model. The method employs a single Transformer denoiser with block-causal attention within a diffusion framework to jointly generate compact, temporally unaligned joint-space path sketches and dense executable action chunks, thereby guiding action generation via geometric intent. Temporal invariance is achieved through arc-length parameterization, while interleaved denoizing preserves the directional dependence of actions on sketches. Extensive evaluations demonstrate that this approach significantly improves success rates by up to 75% in the Push-T simulation and four real-world bimanual tasks, validating its effectiveness.
📝 Abstract
Generative robot policies predict short action chunks but lack explicit long-horizon intent. Recent methods expose longer-horizon structure through language plans, subgoal images, or video forecasts, which are costly to generate and still need to be translated into robot motion. Predicting future robot motions avoids this translation, but a dense, time-indexed trajectory requires numerous parameters to cover the full remaining task, and over a short horizon it largely repeats the action chunk and adds little guidance for action generation. We propose Proprioceptive Action Models (PAM), which jointly generate a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk within a single transformer denoiser. The sketch parameterizes the path by arc length rather than time, capturing geometric intent invariant to execution timing. Block-causal attention and a staggered denoising schedule maintain directed sketch-to-action dependence, ensuring the action tokens condition on a progressively cleaner sketch throughout sampling. In simulation, PAM improves over its action-only counterparts on Push-T and LIBERO-Long; on four real-world bimanual tasks, it raises success from 47.5% to 75.0%. Project page: https://nicehiro.github.io/pam_dp/
Problem

Research questions and friction points this paper is trying to address.

generative robot policies
long-horizon intent
action generation
robot motion prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proprioceptive Action Models
Generative robot policies
Arc-length parameterization
Block-causal attention
Staggered denoising
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Fangyuan Wang
Fangyuan Wang
AMD
S
Songhao Huang
Hong Kong Polytechnic University
H
Haoxiang Sun
Hong Kong Polytechnic University
Shipeng Lyu
Shipeng Lyu
EITs
C
Chengyang He
National University of Singapore
Anqing Duan
Anqing Duan
MBZUAI
Robotics
P
Peng Zhou
Great Bay University
David Navarro-Alarcon
David Navarro-Alarcon
The Hong Kong Polytechnic University
Robotics