Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

๐Ÿ“… 2025-03-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing image quality assessment (IQA) methods suffer from overreliance on opaque numerical scores or large-scale annotated datasets, lacking content awareness and interpretability. To address this, we propose the first vision-oriented reinforcement learning framework based on Group Relative Policy Optimization (GRPO). Our method jointly models score regression and degradation classification, enabling content-aware quality understanding, fine-grained degradation identification, and zero-shot pairwise image comparisonโ€”using only a small number of scalar quality scores and degradation labels. On benchmark IQA tasks, our approach significantly outperforms state-of-the-art methods in both score regression and degradation classification. Crucially, it demonstrates strong generalization to unseen distortions and robust zero-shot comparative reasoning, overcoming key limitations of conventional IQA approaches in flexibility, interpretability, and few-shot adaptability.

Technology Category

Computer Vision: Learning & Optimization for CVIntelligent Robots: Learning & Optimization for ROBKnowledge Representation and Reasoning: Qualitative Reasoning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
๐Ÿ“ Abstract
Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language models (MLLMs) has significantly broadened the scope of IQA, moving toward comprehensive image quality understanding that incorporates content analysis, degradation perception, and comparison reasoning beyond mere numerical scoring. Previous MLLM-based methods typically either generate numerical scores lacking interpretability or heavily rely on supervised fine-tuning (SFT) using large-scale annotated datasets to provide descriptive assessments, limiting their flexibility and applicability. In this paper, we propose Q-Insight, a reinforcement learning-based model built upon group relative policy optimization (GRPO), which demonstrates strong visual reasoning capability for image quality understanding while requiring only a limited amount of rating scores and degradation labels. By jointly optimizing score regression and degradation perception tasks with carefully designed reward functions, our approach effectively exploits their mutual benefits for enhanced performance. Extensive experiments demonstrate that Q-Insight substantially outperforms existing state-of-the-art methods in both score regression and degradation perception tasks, while exhibiting impressive zero-shot generalization to comparison reasoning tasks. Code will be available at https://github.com/lwq20020127/Q-Insight.
Problem

Research questions and friction points this paper is trying to address.

Enhance image quality assessment via visual reinforcement learning
Reduce reliance on large annotated datasets for quality understanding
Improve interpretability and flexibility in image quality evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses reinforcement learning for image quality
Optimizes score regression and degradation tasks
Requires minimal labeled data for training
๐Ÿ”Ž Similar Papers
W
Weiqi Li
School of Electronic and Computer Engineering, Peking University
X
Xuanyu Zhang
School of Electronic and Computer Engineering, Peking University
S
Shijie Zhao
ByteDance Inc.
Y
Yabin Zhang
ByteDance Inc.
Junlin Li
Junlin Li
ByteDance Inc. - Georgia Institute of Technology - Tsinghua University
Video Compression and ProcessingVideo StreamingMachine LearningAIASIC Design
L
Li Zhang
ByteDance Inc.
J
Jian Zhang
School of Electronic and Computer Engineering, Peking University