Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overlooked role of timestep weighting in ELBO-based reinforcement learning and the performance limitations caused by inheriting pretraining configurations. For the first time, timestep weights are treated as a core design element. Through gradient analysis, we reveal that optimal weighting depends on both the reward landscape and the training stage. Accordingly, we propose static, budgeted, and behavior-gap-based dynamic scheduling strategies, providing a unified explanation for existing heuristic loss choices. Experiments on CIFAR generation and robotic tasks demonstrate that our dynamic schemes significantly outperform conventional fixed objectives, establishing the critical importance of timestep weighting in flow matching RL fine-tuning.
📝 Abstract
ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.
Problem

Research questions and friction points this paper is trying to address.

Timestep Weighting
ELBO-based Reinforcement Learning
Flow Matching
Reward Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Timestep Weighting
Flow-Matching RL
ELBO
Gradient Analysis
Dynamic Scheduling