π€ AI Summary
Diffusion-based super-resolution models frequently generate hallucinated details, rendering existing quality assessment metrics ineffective. This study addresses this limitation by constructing DISRQAD, the first diagnostic benchmark tailored for this task, comprising 140,000 subjective scores to systematically evaluate 51 quality metrics. Furthermore, by integrating knowledge distillation with model pruning, this work proposes Q-ReAlign-mini, a lightweight assessment model. Our analysis reveals a substantial misalignment between standard metrics and human perception, with the best-performing no-reference metric achieving a Spearmanβs rank correlation coefficient (SRCC) of only 0.431. In contrast, Q-ReAlign-mini improves the SRCC to 0.496, effectively bridging the evaluation gap in diffusion-based super-resolution scenarios.
π Abstract
Diffusion-based image super-resolution (SR) can create visually plausible detail that is not supported by the low-resolution input. We introduce DISRQAD, a subjective-quality dataset and diagnostic benchmark for this setting. It contains mean opinion scores (MOS) for 14,000 SR outputs from ten diffusion and four non-diffusion methods, spanning four low-resolution degradation conditions and x2/x4 upscaling. We evaluate 51 standard full-reference and no-reference metric configurations and 11 adapted variants. Agreement with MOS is substantially weaker on diffusion outputs: the strongest standard no-reference baseline reaches 0.431 SRCC on diffusion SR versus 0.813 on non-diffusion SR. As a case study in benchmark use, a pruned and distilled Q-ReAlign-mini student reaches 0.496 SRCC on diffusion SR. DISRQAD measures perceived output quality, not faithfulness to the input; it enables analysis of metric behavior across generator families and input conditions. Our findings reveal a substantial gap in the assessment of diffusion-based SR and provide a basis for developing quality models sensitive to diffusion-specific artifacts.