🤖 AI Summary
This study addresses the high inference cost of diffusion Transformers by proposing a training-free acceleration method. The approach evaluates outputs only at anchor steps and derives prediction rules based on chord-tangent transport theory to geometrically interpolate and skip intermediate steps. Furthermore, it establishes a geometric anchor spacing principle that minimizes gap expansion, integrating round-trip linear projection with geometry-preserving scores to safeguard generation quality. Experimental results demonstrate that the proposed method achieves an approximate 5× speedup while enhancing fidelity, yielding a 3.10 dB PSNR improvement on the FLUX model. On HunyuanVideo, it attains a 4.99× acceleration alongside a 5.44 dB PSNR gain, significantly reducing computational overhead while effectively improving generative fidelity.
📝 Abstract
Diffusion transformers incur substantial inference cost through repeated model evaluations along a sampling trajectory. We introduce GeoShrink, a training-free acceleration method that retains the original solver grid while evaluating the model only at a prescribed set of anchors. At skipped stages, GeoShrink predicts the solver-facing output by adding a geometrically retained fraction of the latest observed innovation to the most recent exact output. We derive this rule from chordal tangent transport and round-trip line projection, and establish a geometric anchor-spacing principle that minimizes the largest adjacent gap expansion under fixed coverage and first span. The analysis characterizes the geometric closure and propagation of prediction errors without assuming access to future model outputs. Experiments cover image, video, motion, and audio generation, together with adapted 3D backends. At approximately $5\times$ acceleration, GeoShrink improves FLUX PSNR by 3.10 dB over the strongest listed baseline. On HunyuanVideo, it achieves a reported $4.99\times$ speedup and improves ChronoMagic-Bench-150 PSNR by 5.44 dB over the strongest listed fidelity baseline. Comparisons at fixed evaluation budgets further show substantial gains on motion, audio, music, and 3D generation.