EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

๐Ÿ“… 2026-08-03
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing video variational autoencoders (VAEs) are not optimized for embodied manipulation scenarios, resulting in redundant and poorly controllable latent representations that hinder world model training and robotic action accuracy. This work proposes EmbodiedVAE, the first video VAE specifically designed for embodied manipulation, featuring a compact and controllable latent space. It employs a dual-encoder, single-decoder architecture combined with an asymmetric spatiotemporal compression module to explicitly disentangle robotic arm motion from background dynamics. Furthermore, a consistency constraint based on optimal transport is introduced to enhance temporal coherence of action-related latent variables. The method achieves higher compression rates while preserving high reconstruction fidelity, significantly improving control accuracy in robotic manipulation tasksโ€”yielding an average PSNR gain of 2 dB over current state-of-the-art video VAE approaches.
๐Ÿ“ Abstract
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllable, thereby hindering efficient training of LDMs and precise robotic control. To solve this problem, we present EmbodiedVAE, a novel video VAE that provides compact yet controllable latent representations tailored for the robotic manipulation world models. Specifically, EmbodiedVAE adopts a dual-encoder, single-decoder architecture with an asymmetric spatio-temporal compression module, which automatically disentangles the robot arm's motion from background environment, resulting in overall compactness while providing explicit embodied latent to support fine-grained action control. To further preserve the temporal consistency of learned robotic motion latent, we introduce an optimal-transport-based consistency module that explicitly enforces motion fidelity and inter-frame coherence. Extensive experiments demonstrate that our proposed EmbodiedVAE achieves superior reconstruction quality with high compression rate, while enabling more precise action control in robotic manipulation scenarios with an average of 2dB PSNR improvement over state-of-the-art video VAEs.
Problem

Research questions and friction points this paper is trying to address.

embodied manipulation
latent representation
video VAE
controllability
compactness
Innovation

Methods, ideas, or system contributions that make the work stand out.

EmbodiedVAE
disentangled representation
video VAE
optimal transport
robotic manipulation
๐Ÿ”Ž Similar Papers
2024-04-02IEEE/RJS International Conference on Intelligent RObots and SystemsCitations: 0