🤖 AI Summary
This study addresses the high computational cost of sampling in flow matching models by proposing ReAL, a training-free adaptive sampler. The method reveals an intrinsic relationship between velocity mismatch and candidate step sizes, leveraging a shared lookahead mechanism to dynamically select optimal steps. Furthermore, it performs corrective updates based on the pretrained velocity field and the original noise schedule, advancing generation with only a single network evaluation per step while avoiding additional endpoint computation overhead. Experimental results demonstrate that ReAL achieves a 4.91× speedup on FLUX.1-dev while retaining 97% of the generation quality, and yields a 5.49× acceleration on HunyuanVideo with VBench scores closely approaching those of dense sampling baselines.
📝 Abstract
Flow-matching models generate high-quality images and videos, but repeated neural network evaluations make sampling expensive. Skipping evaluations reduces this cost by extending an available velocity estimate over a longer span. However, local velocity agreement alone does not determine a suitable span, and checking each candidate endpoint adds costly model calls. We introduce ReAL, a training-free sampler that selects how far to advance using one shared lookahead. Our key insight is that the discrepancy between uncorrected and lookahead-corrected endpoint proposals can be computed directly from the observed velocity mismatch and the candidate span beyond the lookahead. This relation provides a span-dependent selection criterion without additional endpoint evaluations. The same lookahead selects the longest passing candidate span, corrects the accepted update, and supplies its velocity as the next starting estimate. After initialization, each regular iteration requires only one fresh evaluation. ReAL uses the pretrained velocity output and original noise schedule, with no additional training or access to internal features. Experiments cover four image-generation backbones, video generation, and image editing. ReAL achieves 4.91x measured speedup on FLUX.1-dev while retaining 97.0% of dense mean ImageReward. On HunyuanVideo, it achieves a 5.49x speedup while maintaining a VBench score close to that of dense sampling.