Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the computational inefficiency and action latency inherent in traditional embodied control methods that rely on real-time, high-cost world-generation models. The authors propose internalizing the multi-level intermediate states produced by a future-prediction generator into a predictive representation trained via supervised learning, which depends solely on current visual–language inputs. This approach enables efficient control without explicit future generation for the first time. By “folding” the generator’s internal structure into a present-moment representation, the method supports rapid adaptation to dynamic environmental perturbations and human interventions. Evaluated on LIBERO, RoboTwin2.0, and real-robot tasks, the framework reduces inference latency by 3.7–10.1×, effectively suppresses irrelevant scene variations, captures long-horizon dynamics, and demonstrates robust flexibility under manual intervention.
📝 Abstract
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by $3.7\times$ relative to Fast--WAM, Enfold-Flash reaches $10.1\times$. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.
Problem

Research questions and friction points this paper is trying to address.

world generative models
embodied control
predictive representations
action latency
efficient inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

predictive representation
world generative model
embodied control
computation folding
latent supervision
🔎 Similar Papers
2024-05-28International Conference on Learning RepresentationsCitations: 10
2024-07-09IEEE/ASME transactions on mechatronicsCitations: 94
W
Weili Zeng
Shanghai Jiao Tong University
Y
Yitong Xing
Shanghai Jiao Tong University
F
Fulong Liu
Shanghai Jiao Tong University
C
Chengqun Yang
Shanghai Jiao Tong University
A
Antao Xiang
South China University of Technology
Feng Tian
Feng Tian
Shanghai Jiao Tong University
Machine Learning
Jingnan Gao
Jingnan Gao
Ph.D. student at Shanghai Jiao Tong University
Computer Vision
J
Jisong Cai
Shanghai Jiao Tong University
X
Xin Wang
Qilu University of Technology (Shandong Academy of Sciences)
X
Xiaomin Wu
Qilu University of Technology (Shandong Academy of Sciences)
Y
Yao Mu
Shanghai Jiao Tong University
Y
Yichao Yan
Shanghai Jiao Tong University