๐ค AI Summary
This study addresses the substantial inference overhead of diffusion world models and the limitation that existing caching methods fail to exploit the underlying mathematical structure of features. To this end, this work proposes a training-free spectral caching framework. Leveraging the temporal stability of feature singular subspaces, the method estimates singular values at subsequent timesteps via linear extrapolation, thereby bypassing redundant backbone network computations and significantly reducing inference costs. Evaluated on the HunyuanWorld-Voyager-13B model, the proposed approach achieves a 5.22ร inference speedup while preserving lossless generation quality. This work establishes a new paradigm for the efficient deployment of large-scale diffusion world models.
๐ Abstract
Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlying mathematical structure of diffusion features largely unexplored. In this work, we reveal that world-model features exhibit highly stable singular subspaces across nearby denoising steps, while their singular values follow predictable evolution patterns. Building on this observation, we propose SpectralCache, a training-free spectral caching framework that reuses stable singular subspaces and estimates only low-dimensional singular values through linear extrapolation. We further exploit the spectral consistency between neighboring full-computation features to skip selected expensive backbone evaluations via singular value scaling. Extensive experiments on representative world models demonstrate that SpectralCache consistently improves inference efficiency while preserving generation quality. On HunyuanWorld-Voyager-13B, SpectralCache achieves 5.22x acceleration while maintaining a WorldScore of 65.90 for static scenes, substantially outperforming existing training-free caching methods in inference efficiency.