OnlineWM: Causality-Aware Active Online Learning for Effective World Modeling

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
针对世界建模中数据分布错位和因果关系缺失问题,提出OnlineWM框架,通过主动在线学习和因果意识微调来提高模型预测的准确性和鲁棒性。
📝 Abstract
Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeling. While recent efforts leverage simulator-generated data to enhance this capability, existing training pipelines face two fundamental limitations. First, static offline data collection leads to a distribution misalignment between training sets and the model's evolving error patterns, failing to resolve critical long-tail scenarios where dynamics predictions remain unreliable. Second, the standard objective of minimizing observational discrepancy often encourages the model to exploit spurious correlations instead of capturing the underlying action-effect causality. To address these limitations, we propose OnlineWM, an online training framework that continuously improves world modeling through active simulator interaction and causality-aware optimization. OnlineWM introduces two key innovations: (1) Active Online Learning: Instead of using fixed datasets, OnlineWM adaptively queries the simulator for new interaction sequences that target the model's current predictive weaknesses, ensuring high-utility data acquisition. (2) Causality-Aware Fine-Tuning: We propose a counterfactual learning strategy that contrasts the outcomes of different actions from identical states, forcing the model to attribute state transitions to specific actions rather than ambient environmental evolution, thereby grounding its predictions in reliable causal mechanisms. By integrating active data acquisition with causal optimization, OnlineWM establishes a closed-loop refinement process that ensures the model is both robust to diverse scenarios and precise in its causal attribution. Extensive experiments demonstrate that OnlineWM significantly enhances action controllability and generalizes effectively to unseen domains.
Problem

Research questions and friction points this paper is trying to address.

generative world models
distribution misalignment
causal mechanisms
action controllability
long-tail scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Online Learning
Causality-Aware Fine-Tuning
World Modeling
🔎 Similar Papers
No similar papers found.
Y
Yikun Miao
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China
F
Fangqi Zhu
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China
Q
Quanxin Shou
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China
X
Xiaoyi Pang
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China
Z
Zhengyang Yan
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China
Junhao Li
Junhao Li
Assistant Project Scientist, Cognitive Science, University of California, San Diego
Non-coding RNAsDNA methylationEpigeneticsBioinformatics
H
Haodong Wang
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China
Zicong Hong
Zicong Hong
Department of Computer Science and Engineering, Hong Kong University of Science and Technology
BlockchainML SystemEdge/Cloud Computing
Song Guo
Song Guo
Chair Professor of CSE, HKUST
Large Language ModelEdge AIMachine Learning Systems