Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses two key limitations of existing world action models: the absence of cross-task skill reuse mechanisms and background redundancy that interferes with critical visual information extraction. To this end, we propose a skill reuse framework based on an action experience dictionary and shared action embeddings. Specifically, a pretrained action tokenizer encodes historical trajectories, while a cross-attention mechanism integrates visual conditions with action intent to enable precise prediction. Furthermore, a motion-aware transition loss is introduced to effectively suppress background noise. Extensive evaluations demonstrate the superiority and generalization capability of the proposed method across both simulation benchmarks and real-world cross-embodiment scenarios.
πŸ“ Abstract
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The anonymous project website is available at \href{https://github.com/JiahuaDong/AED}{AED}.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
skill reuse
cross-task semantic relationships
action experience
visual dynamics prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action Experience Dictionary
World Action Models
Skill Reuse
Motion-aware Transition Loss
Cross-embodiment
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Qi Lyu
Qi Lyu
Master of Science, Michigan State University
Deep LearningNLP
J
Jiahua Dong
Mohamed bin Zayed University of Artificial Intelligence
H
Hao Shen
Anhui University
X
Xudong Wang
Shenyang Institute of Automation, Chinese Academy of Sciences
H
Hongyuan Yu
Xiaomi Corporation
B
Baichen Liu
Shenyang Institute of Automation, Chinese Academy of Sciences
Henghui Ding
Henghui Ding
Fudan University
Computer VisionMachine LearningSegmentationAIGC
Zhi Han
Zhi Han
SIA, CAS
Computer Vision
Nicu Sebe
Nicu Sebe
University of Trento
computer visionmultimedia
Ivan Laptev
Ivan Laptev
Professor at MBZUAI, on leave from INRIA
Computer VisionRoboticsAction RecognitionObject Recognition
Fahad Shahbaz Khan
Fahad Shahbaz Khan
MBZUAI, LinkΓΆping University Sweden
Computer VisionObject RecognitionGenerative AIAI for Science
S
Salman Khan
Mohamed bin Zayed University of Artificial Intelligence