G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the negative transfer problem in video model post-training, where heterogeneous reward data induces skill conflicts that degrade joint training performance. To mitigate this issue, we propose a gradient compatibility-based data organization strategy. Specifically, heterogeneous data is first partitioned into distinct groups via gradient probing and clustering, with dedicated LoRA experts trained for each group. Subsequently, these multiple experts are consolidated into a single adapter through weight merging and online policy distillation. This approach effectively alleviates negative transfer among diverse skills. Extensive experiments demonstrate that our method significantly improves VBench scores on models such as Wan2.1, consistently outperforming existing baseline approaches.
📝 Abstract
Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2.0 dimensions define data buckets and an external multimodal reward pipeline assigns sample weights. We propose G$^3$-LoRA (Gradient-Guided Grouped LoRA), a data organization procedure that probes category-level gradients induced by reward-weighted video samples, removes the shared global update direction, clusters categories by residual gradient compatibility, trains group-specific LoRA experts, and consolidates them into one adapter by weight merging followed by on-policy distillation from the experts. We motivate this procedure by viewing reward-weighted flow matching as velocity-field regression: incompatible reward dimensions may prefer different denoising directions in overlapping noisy latent regions, causing shared LoRA training to average capabilities. On Wan2.1-T2V-1.3B-Diffusers, the merged grouped adapter improves the matched VBench2.0 evaluation over the base model, a joint reward-weighted LoRA baseline, and random, semantic, and raw-gradient partitions trained with the same pipeline; an independent evaluator agrees, and on CogVideoX-2B grouping avoids the negative transfer of joint training. The gain is not uniform: merging compresses the largest specialist gains, distillation recovers part of this loss, and camera motion and several local-quality dimensions remain challenging. Together, these results suggest that gradient compatibility can serve as a practical diagnostic for organizing reward-weighted video post-training data.
Problem

Research questions and friction points this paper is trying to address.

reward-weighted video data
text-to-video post-training
gradient compatibility
negative transfer
heterogeneous data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient-Guided Grouped LoRA
Reward-Weighted Video Post-Training
Residual Gradient Compatibility
Weight Merging and Distillation
Flow Matching
🔎 Similar Papers
No similar papers found.