PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent contradiction between the prohibitive computational overhead of diffusion models and the limited detail recovery capability of lightweight networks, proposing a generative heterogeneous distillation framework for real-world image super-resolution with zero additional inference cost. Methodologically, the approach transfers diffusion priors into a feed-forward network via score-based distribution matching. Furthermore, a directional reliability weighting mechanism is designed to circumvent inefficient feature alignment or output imitation, thereby achieving high-fidelity knowledge transfer. Extensive experiments demonstrate that the proposed method significantly enhances perceptual quality while preserving reconstruction fidelity across multiple benchmarks and backbone architectures, all without introducing any extra inference overhead.
📝 Abstract
Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Transformer SR models are efficient yet often struggle to recover realistic high-frequency details. This motivates a natural question: can diffusion priors be transferred to existing diffusion-free SR networks without introducing diffusion components at inference time? To this end, we propose PhoenixSR, a generative heterogeneous distillation framework that transfers diffusion priors to independently designed feed-forward SR networks through score-based distribution matching. Rather than aligning heterogeneous features or imitating sampled diffusion outputs, PhoenixSR uses the pretrained diffusion model as distribution-level supervision, while paired SR supervision preserves reconstruction fidelity. To make distribution matching effective for fidelity-sensitive SR, we introduce Heterogeneous Distribution Adaptation, which adapts the target score to the SR domain, improves tracking of the evolving student distribution, and anchors training with paired supervision. We further employ Directional Reliability Weighting, a lightweight residual-consistency-based reweighting strategy that reduces unstable distributional guidance. All diffusion-related components are removed after training, leaving the original student architecture and inference cost unchanged. Experiments on three SR benchmarks and six feed-forward backbones, including SwinIR, HAT, Real-ESRGAN, and SeeMoRe, show consistent perceptual improvements with largely preserved reconstruction fidelity.
Problem

Research questions and friction points this paper is trying to address.

Real-World Super-Resolution
Diffusion Prior
Knowledge Distillation
Perceptual Quality
Computational Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Heterogeneous Distillation
Score-based Distribution Matching
Real-World Super-Resolution
Heterogeneous Distribution Adaptation
Directional Reliability Weighting