Time dependent loss reweighting for flow matching and diffusion models is theoretically justified

📅 2025-11-20
📈 Citations: 0
Influential: 0
📄 PDF

career value

192K/year
🤖 AI Summary
This work addresses loss function design within the generator matching framework—encompassing flow matching and diffusion models over continuous, manifold, and discrete spaces—and identifies key limitations in existing formulations. Method: Leveraging stochastic process modeling, variational inference, and convex analysis, we unify diverse generative dynamics under a principled framework. We prove that (1) Bregman divergence losses and linearly parameterized generators can explicitly depend on both state $X_t$ and time $t$, and (2) temporal expectations may be taken with respect to generalized time distributions—not merely uniform ones—thereby justifying common time-weighting heuristics. We extend these insights to Edit Flows, revealing structural equivalence in loss design. Contribution/Results: Our analysis broadens the theoretical scope of generator matching, simplifies the construction of $X_1$-predictors, and demonstrates consistent empirical gains across multiple flow- and diffusion-based models.

Technology Category

Application Category

📝 Abstract
This brief note clarifies that, in Generator Matching (which subsumes a large family of flow matching and diffusion models over continuous, manifold, and discrete spaces), both the Bregman divergence loss and the linear parameterization of the generator can depend on both the current state $X_t$ and the time $t$, and we show that the expectation over time in the loss can be taken with respect to a broad class of time distributions. We also show this for Edit Flows, which falls outside of Generator Matching. That the loss can depend on $t$ clarifies that time-dependent loss weighting schemes, often used in practice to stabilize training, are theoretically justified when the specific flow or diffusion scheme is a special case of Generator Matching (or Edit Flows). It also often simplifies the construction of $X_1$-predictor schemes, which are sometimes preferred for model-related reasons. We show examples that rely upon the dependence of linear parameterizations, and of the Bregman divergence loss, on $t$ and $X_t$.
Problem

Research questions and friction points this paper is trying to address.

Theoretically justifies time-dependent loss weighting in flow matching models
Extends loss function flexibility to depend on both state and time variables
Simplifies construction of predictor schemes for diffusion model training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time-dependent loss reweighting stabilizes training theoretically
Bregman divergence loss depends on both state and time
Linear generator parameterization varies with time and state