FIRM-WM: State-factorized factual-interventional recurrent modeling for reward-free visual planning

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出FIRM-WM模型,通过分离目标可比配置与动态预测纤维解决无奖励视觉规划中的状态表示问题,并在多个任务上实现更高效能。
📝 Abstract
Reward-free latent world models can learn from offline videos and solve new image--goal tasks by optimizing actions through predicted latent futures. This setting places two demands on the planning state: its coordinates must be comparable with a goal image. Moreover, its dynamics must retain velocity, motion trend, contact, and other history--dependent information beyond those goal coordinates. Offline training creates a second mismatch: each recorded trajectory reveals one factual future, whereas a sampling--based planner compares many actions that were not taken from the same state. We introduce FIRM-WM (Factual--Interventional Recurrent World Model), a compact pixel world model designed around these two gaps. Its recurrent state separates a typed, goal--comparable configuration from a 128-dimensional dynamic fiber used for prediction but excluded from the terminal goal cost. Broad factual trajectories provide state coverage, while common--reset intervention branches provide observed outcomes for alternative action sequences. Before executing each branch, we reset the environment and restore the same recorded values exposed by the environment's state--setting interface. Under matched CEM planning and three independent full-pipeline seeds, FIRM-WM reaches 99.0$\pm$1.0% on TwoRoom, 92.7$\pm$2.1% on Reacher, and 88.0$\pm$3.0% on OGBench-Cube, compared with 89.0%, 88.0%, and 70.0% for LeWM. The deployed model uses 2.98--3.42M parameters and records 2.13--11.60$\times$ lower planning time on these tasks.
Problem

Research questions and friction points this paper is trying to address.

reward-free visual planning
latent world models
offline training
factual trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

state-factorized
factual-interventional
recurrent world model
reward-free visual planning
latent dynamics
Y
Yilun Wu
Saint Petersburg State University, Saint Petersburg, Russia
Y
Yunjian Zhang
Shenzhen Kaihong Digital Industry Development Co., Ltd., Shenzhen, China
Aobo Li
Aobo Li
Assistant Professor, University of California San Diego
Neutrino PhysicsMachine Learning
M
Mujiangshan Wang
Chinese Academy of Sciences (CAS) – Shenzhen Institute of Advanced Technology; Shenzhen Kaihong Digital Industry Development Co., Ltd., Shenzhen, China
Haitao Wu
Haitao Wu
Microsoft
NetworkingDatacenterQoSTCP/IPWireless
A
Aqiang Zhang
Harbin Institute of Technology, Harbin, China; Chongqing Research Institute of HIT, Chongqing, China