Score
Designs and implements training objectives, loss functions, and supervision procedures that match continuous flow/vector fields in a learned perceptual feature space rather than raw input space. This involves constructing the perceptual feature mapping and flow-matching targets so the model’s dynamics align with perceptual distances, bias regression toward on-manifold modes, enable high-quality few-step sampling, and avoid reliance on teacher or score networks.
This work addresses the limitations of existing few-step generation methods, which often rely on teacher models or auxiliary networks, thereby compromising efficiency and generalization. The authors propose Perceptual Flow Matching (PFM), a novel framework that, for the first time, transfers flow matching supervision from the conventional VAE latent space to a pretrained perceptual feature space, replacing standard velocity regression. PFM requires neither additional networks nor knowledge distillation; instead, it achieves highly efficient few-step synthesis solely through a shift in representation space. The approach reveals that perceptual supervision effectively guides optimization toward authentic data manifold modes. Evaluated across image and video generation as well as image editing tasks, PFM reduces sampling steps from 35–50 to just 4–8 while maintaining high visual quality and significantly suppressing artifacts, outperforming current distillation-based methods.
This work addresses the limitations of current large vision-language models, which suffer from language bias and hallucination due to optimization objectives that inadequately constrain visual grounding, and rely on geometric priors that prioritize precision over reasoning utility. To overcome these issues, the authors propose Perception Flow Network (PFlowNet), a framework that decouples perception and reasoning to establish a self-conditioned generation process. PFlowNet eschews rigid alignment with expert priors and instead employs variational reinforcement learning to drive perception toward reasoning goals. It further integrates a multi-dimensional reward mechanism with neighborhood geometric shaping to enhance both interpretability and reasoning efficiency. The model achieves state-of-the-art performance at the time of evaluation, attaining 90.6% on V* Bench and 67.0% on MME-RealWorld-lite.
This paper identifies the fundamental cause of topological discontinuities in velocity fields within flow matching: when the prior (e.g., unimodal) and target distribution (e.g., multimodal) exhibit topological mismatch, the optimal velocity field necessarily develops asymptotically infinite jump discontinuities along decision boundaries. This arises from the geometric necessity for continuous flows to split particle trajectories to map distinct modes—not from loss function design or optimization bias. Method: We adopt a topological dynamical systems perspective, providing theoretical analysis for a bimodal Gaussian mixture, embedding the problem within a Riemannian flow matching framework, and validating findings via neural representation experiments. Contribution/Results: We rigorously establish that such discontinuities are ubiquitous at intermediate times along decision boundaries. The phenomenon is model-agnostic—persisting across diverse loss functions and flow architectures—and imposes critical theoretical constraints on extending flow matching to nontrivial manifolds.
This work addresses the theoretical gap in sample complexity analysis for flow-matching generative models. Unlike prior studies relying on empirical risk minimization (ERM) assumptions, we establish the first end-to-end upper bound on sample complexity without such assumptions. Methodologically, we model the continuous flow via ordinary differential equations and parameterize the velocity field using neural networks; we then introduce a triple-error decomposition framework—comprising neural approximation error, statistical error, and optimization error—and rigorously analyze its convergence. Our theoretical analysis shows that $O(varepsilon^{-4})$ samples suffice to achieve $O(varepsilon)$ generative accuracy in the Wasserstein-2 distance. This constitutes the first rigorous, non-ERM-dependent sample complexity guarantee for flow matching, filling a critical theoretical void. Moreover, our result provides foundational insights for efficient training and generalization analysis of flow-based generative models.
This work addresses the fundamental challenge in transfer learning: the difficulty of accurately characterizing task similarity to reliably predict transfer performance. We propose a novel theoretical framework grounded in feature-space overlap—not distributional distance—enabling principled transfer prediction. Through analytical modeling and phase-transition analysis of deep linear networks, we rigorously establish feature representation consistency as the necessary and sufficient condition for successful transfer, and construct an analytically tractable transfer phase diagram. Our theory reveals a sharp phase transition in transfer performance governed jointly by source-task sample size and feature overlap. Empirically, under high overlap, linear transfer and fine-tuning significantly outperform training from scratch; these findings are validated numerically on nonlinear networks. Crucially, this work overturns the conventional paradigm that relies on φ-divergences or other distributional metrics for transfer assessment, providing both a rigorous theoretical foundation and practical criteria for few-shot transfer learning.
This work establishes the first rigorous theoretical foundation for neural network–based flow matching, addressing the lack of guarantees regarding convergence, generalization, and generation quality. Focusing on over-parameterized two-layer ReLU neural networks that model conditional velocity fields, the study analyzes the Wasserstein error of samples generated by the induced flow under gradient descent optimization and provides the first convergence and generalization bounds for flow matching. Furthermore, it introduces a novel generalization theory for multi-task representation learning applicable to unbounded losses. Empirical evaluations on both synthetic data and real-world image benchmarks corroborate the theoretical predictions, demonstrating that the method achieves strong convergence properties alongside high-quality sample generation.
Existing visual world models struggle to preserve information useful for downstream perception tasks while modeling future uncertainty. This work introduces stochastic flow matching in high-dimensional pretrained visual feature spaces—such as DINOv3—for the first time, constructing a stochastic world model equipped with a differentiable single-step projection mechanism tailored to this space to enable efficient training. By integrating temporal consistency constraints with task-driven optimization objectives, the proposed method substantially improves performance on perception tasks, enhances mode coverage in multimodal future prediction, and increases robustness in long-horizon forecasting across both synthetic and real-world benchmarks, demonstrating its effectiveness and generalizability.
This work addresses the limitations of existing flow matching methods, which rely on fixed mean squared error losses and struggle to accurately align with target data distributions. The authors propose Continuous Adversarial Flow (CAF), the first framework to integrate adversarial training into continuous-time flow matching by replacing conventional loss functions with a learnable discriminator that dynamically guides the generation process. CAF supports both end-to-end training and serves as a general-purpose post-optimization strategy to enhance pre-trained flow matching models. Evaluated on ImageNet at 256px resolution, CAF achieves state-of-the-art unconditional FID scores of 3.63 (with SiT) and 3.57 (with JiT), while also improving performance in guided image generation and text-to-image tasks, demonstrating its effectiveness in enhancing sample quality and distribution alignment.
This study addresses the limitations of flow matching, which is constrained by fixed interpolation paths and prone to path overfitting during joint training. To overcome these issues, this work proposes a path-flow alignment loss that jointly optimizes an endpoint-preserving path network and a flow network to enable their co-evolution. Furthermore, a stochastic path regularizer is designed to introduce an explicit entropy lower bound, thereby suppressing low-entropy bottlenecks and ensuring training stability. The proposed approach is compatible with model-guided training and requires no modifications to the inference architecture. Experiments on ImageNet demonstrate significant improvements in FID scores, establishing a more flexible and efficient paradigm for path optimization within flow matching frameworks.
This work addresses the inherent mismatch between continuous flow matching and discrete semantic segmentation tasks, which leads to vanishing gradients, trajectory crossing, slow convergence, and blurred class boundaries. From the perspective of vector field learning, the authors propose reshaping the velocity field and introducing a distance-aware correction term to enhance inter-class separability. Additionally, they design a class encoding scheme inspired by quasi-random Kronecker sequences to enable efficient gradient propagation and pixel-level semantic alignment under end-to-end training. The proposed method significantly outperforms the original flow-matching approach and substantially narrows the performance gap between generative segmentation models and strong discriminative counterparts across multiple metrics.