🤖 AI Summary
Diffusion models face dual challenges of high computational overhead and severe hardware resource constraints when deployed on edge devices. This paper proposes the first full-stack optimization framework for diffusion models tailored to edge scenarios, systematically integrating lightweight architectural design, adaptive denoising step scheduling, mixed-precision quantization, structured pruning, knowledge distillation, heterogeneous compute offloading, and a customized inference engine. We innovatively establish an edge-adapted sampling acceleration paradigm and lightweight architecture design principles, bridging the gap between theoretical optimization and practical edge deployment. Evaluated on edge platforms such as the Jetson Orin, our framework achieves a 5.3× speedup in inference latency and a 78% reduction in memory footprint for 1024×1024 image generation—enabling, for the first time, real-time, high-fidelity cross-modal generation under stringent edge resource constraints.
📝 Abstract
Diffusion models have shown remarkable capabilities in generating high-fidelity data across modalities such as images, audio, and video. However, their computational intensity makes deployment on edge devices a significant challenge. This survey explores the foundational concepts of diffusion models, identifies key constraints of edge platforms, and synthesizes recent advancements in model compression, sampling efficiency, and hardware-software co-design to make diffusion models viable on edge devices. We also review promising applications and suggest future research directions.