🤖 AI Summary
This study addresses whether training losses for generative flows guarantee sampling accuracy, noting that theoretical bounds on loss convergence to zero and gradient descent rates remain unclear. By integrating total variation distance with the spectral theory of Green's operators, this work comparatively analyzes error bounds between balance loss and flow matching loss, with all proofs formally verified in Lean 4. The results reveal that ratio-based losses under continuous generators cannot bound total variation error. Furthermore, it is proven that squared logarithmic generators achieve global convergence from arbitrary positive initializations on finite connected graphs. By deriving curvature constants based on Green's operator norms, this paper establishes explicit error bounds and lower bounds on convergence rates, with all theorems mechanically certified using Lean 4.
📝 Abstract
Generative flows sample from an unnormalized target by training a flow to be balanced, and the training loss is the signal a practitioner watches. We ask what that signal is worth: whether a small loss certifies an accurate sampler, whether the loss can be driven to zero, and how fast gradient descent does so. The loss decides the first. Losses that compare the two sides of the balance by their difference bound, in total variation, the error of the sampler the flow implies, with explicit constants that do not involve the policy; flow-matching losses that compare them through a ratio admit no such bound, already on a single cycle, whenever their generator is continuous at balance. On graphs, the backward policy decides the other two. Once it is frozen, balance becomes invariance under the backward chain, so that existence is free on finite graphs, and one constant --- the norm of that chain's Green operator, which plays the role of an inverse spectral gap --- fixes the order of the curvature of the loss around the balanced flow, from above and below, and sets a floor under the rate at which training converges near it. The mechanism is that gradient descent diffuses the flow along the backward policy. For the squared-logarithm generator of detailed and trajectory balance, training the balance loss on states converges globally on every finite path-connected graph, from every positive initialization. The constant can be infinite while backward trajectories are short on average, and exact flow matching can then fail. The bounds and rates are tested by exact computation on enumerable state spaces, and every theorem carries a certification status computed from a Lean~4 development.