🤖 AI Summary
This study addresses the high decoding latency and real-time planning challenges caused by causal Conv3D operations when deploying generative world models on edge devices. To this end, we propose a cache-aware degradation strategy that requires no model modifications. By leveraging compilation techniques to transform causal Conv3D into batched spatial Conv2D, this method achieves lightweight runtime optimization while fully preserving pretrained weights and semantic representations, thereby eliminating the need for fallback mechanisms. Experiments conducted on the Jetson AGX Orin platform using the Cosmos3 and LingBot-World architectures demonstrate an approximately 7× acceleration in VAE decoding and over a 2× reduction in overall generation latency. The proposed approach attains performance comparable to TensorRT with significantly lighter deployment overhead, exhibiting strong cross-model transferability.
📝 Abstract
Generative world models can provide visual rollouts for embodied planning, yet their feasibility on edge devices depends not only on the learned model but also on how the execution runtime represents its operations. We introduce a cache-aware lowering that expresses supported causal Conv3D calls as batched spatial Conv2D operations while preserving pretrained weights, temporal-cache semantics, convolution parameters, bias placement, and output layout. Across the complete Cosmos3-Edge image-to-video pipeline on a 64-GB NVIDIA Jetson AGX Orin, the proposed route accelerates VAE decoding by approximately $7\times$ and reduces complete-generation latency by more than $2\times$, while repeated decoder evaluations maintain complete fast-path coverage without fallbacks. The unchanged lowering also improves Cosmos3-Nano and transfers to LingBot-World's architecturally distinct Wan2.1 VAE. A clean-device comparison against fully specialized TensorRT shows that TensorRT provides a further $1.36\times$ steady-state improvement, but requires substantially greater per-module and per-runtime-state AOT specialization. Same-latent BF16 and FP32 evaluations characterize the finite-precision differences introduced by the alternative execution order. Together, these results position cache-aware lowering as a lightweight runtime optimization that recovers most of the available decoder acceleration without modifying the learned models themselves.