๐ค AI Summary
Existing image quality assessment (IQA) methods suffer from overreliance on opaque numerical scores or large-scale annotated datasets, lacking content awareness and interpretability. To address this, we propose the first vision-oriented reinforcement learning framework based on Group Relative Policy Optimization (GRPO). Our method jointly models score regression and degradation classification, enabling content-aware quality understanding, fine-grained degradation identification, and zero-shot pairwise image comparisonโusing only a small number of scalar quality scores and degradation labels. On benchmark IQA tasks, our approach significantly outperforms state-of-the-art methods in both score regression and degradation classification. Crucially, it demonstrates strong generalization to unseen distortions and robust zero-shot comparative reasoning, overcoming key limitations of conventional IQA approaches in flexibility, interpretability, and few-shot adaptability.
๐ Abstract
Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language models (MLLMs) has significantly broadened the scope of IQA, moving toward comprehensive image quality understanding that incorporates content analysis, degradation perception, and comparison reasoning beyond mere numerical scoring. Previous MLLM-based methods typically either generate numerical scores lacking interpretability or heavily rely on supervised fine-tuning (SFT) using large-scale annotated datasets to provide descriptive assessments, limiting their flexibility and applicability. In this paper, we propose Q-Insight, a reinforcement learning-based model built upon group relative policy optimization (GRPO), which demonstrates strong visual reasoning capability for image quality understanding while requiring only a limited amount of rating scores and degradation labels. By jointly optimizing score regression and degradation perception tasks with carefully designed reward functions, our approach effectively exploits their mutual benefits for enhanced performance. Extensive experiments demonstrate that Q-Insight substantially outperforms existing state-of-the-art methods in both score regression and degradation perception tasks, while exhibiting impressive zero-shot generalization to comparison reasoning tasks. Code will be available at https://github.com/lwq20020127/Q-Insight.