🤖 AI Summary
This study addresses the fundamental contradiction between the stochastic nature of diffusion-based video generation and the determinism required for visual reasoning, which constrains zero-shot inference capabilities. To overcome this limitation, we propose a self-consistency mechanism that aggregates multi-path predictions, coupled with Rejection Fine-Tuning (RFT) to distill the resulting consensus back into the model. Furthermore, training-free test-time scaling and early readout strategies are designed to internalize the advantages of multi-sampling while reducing inference overhead. Without relying on ground-truth supervision, the proposed approach elevates visual search accuracy from 48.4% to 99.0% and achieves an 84.0% success rate in maze solving, thereby enabling efficient and precise perceptual reasoning within a single generation pass.
📝 Abstract
Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.