🤖 AI Summary
This study addresses the misalignment between likelihood-driven objectives and task-specific discriminative boundaries in diffusion model distillation, alongside insufficient geometric coverage of the latent space. To this end, it proposes the MGPO framework, which reformulates data distillation as a multi-objective reinforcement learning problem. By integrating pixel-space discriminative rewards with minimum spanning tree (MST)-based latent geometric rewards, MGPO achieves dual-space alignment that simultaneously preserves inter-class separability and manifold diversity coverage, supported by rigorous theoretical bounds. Furthermore, the framework supports modular extension to structured tasks such as object detection. Under low computational budgets, it yields a substantial 8.0% mIoU improvement on segmentation benchmarks, effectively mitigating sample mode collapse and comprehensively outperforming existing methods.
📝 Abstract
Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage loss, where generated samples collapse into a few high-density modes and fail to cover the manifold's structural diversity. We propose Manifold-Guided Policy Optimization (MGPO), which reformulates DD as a multi-objective reinforcement learning problem and achieves Dual-Space Alignment via a pixel-space discriminative reward and a latent-space geometric reward guided by a class-wise Minimum Spanning Tree (MST). The discriminative reward enforces class separability, while the MST-based geometric reward encourages generated latents to cover a sparse geometric skeleton of each class, jointly addressing both failure modes. We further provide an idealized analysis that motivates the MST-based reward, including a Hausdorff approximation bound and a subsampling bound independent of the dataset size. The reward-modular design extends to structured tasks such as object detection and segmentation by substituting the frozen task reward model. Extensive experiments show MGPO consistently outperforms existing methods, including a +8.0% mIoU gain on segmentation under low-budget settings.