Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high deployment latency and memory constraints of the MiniMax-H3 video diffusion model, which stem from its massive parameter count and multi-step denoising process, by constructing a full-stack inference acceleration pipeline. Methodologically, it proposes a cross-resolution two-stage scheduling strategy coupled with latent-space mapping to eliminate redundant VAE encoding and decoding operations. Furthermore, a recursive self-improvement (RSI) mechanism is introduced to automatically search for optimal kernel fusion and memory layout configurations. Experimental results demonstrate that this approach achieves up to 30Γ— end-to-end speedup and a 20% reduction in GPU memory consumption. Notably, a single DGX Spark card can complete video generation within one minute, effectively overcoming the computational bottlenecks associated with cloud-based deployment.
πŸ“ Abstract
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.
Problem

Research questions and friction points this paper is trying to address.

Video Diffusion Models
Inference Acceleration
Computational Overhead
Generation Latency
Memory Constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Diffusion Models
Cross-Resolution Scheduler
Latent-to-Latent Mapping
Recursive Self-Improvement
Inference Acceleration
πŸ”Ž Similar Papers
No similar papers found.