Score
Design and implement reinforcement-learning systems that produce image quality predictions by treating quality-score selection or ranking as policy actions and optimizing policies using rewards derived from quality errors or correlation-based objectives; construct state representations, action spaces, reward functions (e.g., margin errors or proxies for PLCC/SRCC), and training procedures so the agent learns quality structure through interactions. Analyze policy behavior, convergence, and how reward and optimization choices affect blind/image quality assessment performance and rank/score correlations.
Existing image quality assessment (IQA) methods suffer from overreliance on opaque numerical scores or large-scale annotated datasets, lacking content awareness and interpretability. To address this, we propose the first vision-oriented reinforcement learning framework based on Group Relative Policy Optimization (GRPO). Our method jointly models score regression and degradation classification, enabling content-aware quality understanding, fine-grained degradation identification, and zero-shot pairwise image comparison—using only a small number of scalar quality scores and degradation labels. On benchmark IQA tasks, our approach significantly outperforms state-of-the-art methods in both score regression and degradation classification. Crucially, it demonstrates strong generalization to unseen distortions and robust zero-shot comparative reasoning, overcoming key limitations of conventional IQA approaches in flexibility, interpretability, and few-shot adaptability.
Existing visual quality assessment (VQA) methods suffer from shallow reasoning, poor score calibration, and weak cross-domain generalization due to reliance on supervised fine-tuning or single-objective ranking. This paper proposes a preference–response decoupled reinforcement learning framework featuring a dual-branch reward mechanism and grouped relative policy optimization, unifying absolute scoring and relative ranking objectives for fine-grained, interpretable quality inference. Technically, it integrates reinforcement fine-tuning of multimodal large language models with spatiotemporal data stream modeling. Evaluated on 10 image quality assessment (IQA) and 5 VQA benchmarks, the method achieves state-of-the-art performance: +5.30% Spearman rank correlation coefficient (SRCC) and +2.15% Pearson linear correlation coefficient (PLCC) on IQA tasks, with significantly improved cross-domain generalization and human-aligned reasoning consistency.
Existing reinforcement learning (RL)-based reasoning-style image quality assessment (IQA) models exhibit strong generalization but suffer from opaque mechanisms, high inference energy consumption, and substantial latency. Method: This paper reveals that their generalization stems from mapping visual representations into compact, cross-domain aligned textual embeddings. Building on this insight, we propose RALI—a lightweight framework that directly aligns image features with such textual representations via contrastive learning, eliminating the need for large language model (LLM) invocation or explicit reasoning steps. RALI integrates RL-guided representation learning, contrastive alignment, and multimodal large model (MLLM)-derived visual features. Contribution/Results: Experiments demonstrate that RALI achieves state-of-the-art generalization performance on IQA benchmarks—on par with leading reasoning-based models—while reducing parameter count and inference time to less than 5% of theirs, significantly enhancing practical deployability.
This work addresses a critical limitation in existing reinforcement learning–based image quality assessment methods, which uniformly update sample weights and thereby amplify noise while neglecting the model’s genuine perceptual sensitivity to image content. To overcome this, we propose Q-Hawkeye, a novel framework that dynamically adjusts sample update weights based on predictive uncertainty and introduces an implicit perceptual loss to strengthen the model’s reliance on authentic visual evidence through original–distorted image pairs. By integrating multi-rollout variance estimation, uncertainty-aware advantage weighting, and policy optimization, Q-Hawkeye achieves significant performance gains over state-of-the-art methods across multiple datasets, demonstrating improved assessment accuracy and enhanced cross-dataset generalization.
Reinforcement learning (RL) is difficult to integrate effectively into diffusion-based image restoration models, as these models prioritize fidelity over generative diversity. Method: This paper proposes DiffRL, a novel training framework that synergistically combines RL with diffusion models. Its core innovations are: (1) a multimodal large language model (MLLM)-driven image quality assessment (IQA) module serving as a differentiable reward proxy; and (2) a difficulty-adaptive weighting mechanism that dynamically balances supervised fine-tuning and RL optimization, enabling coarse-to-fine progressive training. The framework is plug-and-play and requires no architectural modification to the underlying diffusion model. Contribution/Results: DiffRL achieves significant performance gains across multiple image restoration tasks—including deblurring, denoising, and super-resolution—particularly demonstrating enhanced robustness on complex or low-quality samples. Extensive experiments validate its effectiveness and strong generalization capability.
This work addresses the limitations of the Qwen-Image-2.0 diffusion model in visual quality and instruction following for text-to-image generation and image editing tasks by proposing an optimization framework that integrates reinforcement learning with online policy distillation. The approach employs a task-oriented composite reward system, an online policy distillation mechanism based on trajectory velocity matching, and combines GRPO algorithm, hybrid classifier-free guidance, and class-level reward calibration to unify multitask capabilities while preserving pretrained knowledge. Experimental results demonstrate significant improvements: the model achieves a total score of 57.84 (+2.61) on Qwen-Image-Bench, with Elo scores rising to 1193 (+78) for text-to-image generation and 1349 (+93) for image editing, reflecting notable gains in aesthetic quality, prompt adherence, and editing accuracy.
This work addresses the challenge of sparse reward signals in reinforcement learning for graphical user interface (GUI) tasks, where success hinges on visual judgment and is difficult to capture via handcrafted rules or human annotations. The authors propose a reinforcement learning fine-tuning framework that leverages autonomous vision-language evaluation, using the alignment score between the final screenshot and the task instruction—generated by a vision-language model—as a terminal reward without requiring task-specific rules or human labels. To mitigate the inherent noise in this reward signal, they introduce a noise-corrected reward estimator integrated with the Proximal Policy Optimization (PPO) algorithm for unsupervised reward generation and policy optimization. Experiments demonstrate that the method improves average success rates by 12.6 percentage points over zero-shot baselines and by 5.1 percentage points over uncorrected reward variants across macOSWorld, Windows Agent Arena, and OSWorld benchmarks.
This work addresses the challenge that autoregressive image generation models trained via maximum likelihood often struggle to balance sample quality and diversity, while existing reinforcement learning approaches are prone to mode collapse. The authors formulate the generation process as a Markov decision process and propose a policy fine-tuning framework based on Group Relative Policy Optimization (GRPO), integrating both instance-level (e.g., CLIP, HPSv2) and distribution-level reward signals. A key innovation is the introduction of leave-one-out FID (LOO-FID) as a distribution-level reward, coupled with explicit diversity promotion through exponential moving average feature moments. Adaptive entropy regularization further enables stable multi-objective optimization. Remarkably, with only a few hundred fine-tuning iterations and without Classifier-Free Guidance, the method significantly enhances both generation quality and diversity while reducing inference cost by approximately half.
This work addresses the limitation of existing image quality assessment (IQA) methods, which typically predict only a single scalar score while neglecting the multidimensional attributes—such as sharpness, color fidelity, noise, and composition—that underpin human perception. To overcome this, the authors propose MG-IQA, a novel framework that jointly predicts overall quality and fine-grained perceptual attributes in a single inference pass. MG-IQA integrates vision-language models with reinforcement learning–based ranking, leveraging attribute-aware prompts, a multidimensional Thurstone reward model, and a cross-domain alignment mechanism to enable interpretable, multi-granularity evaluation without requiring perceptual scale realignment. Experiments demonstrate that MG-IQA consistently outperforms state-of-the-art methods across eight IQA benchmarks, achieving an average 2.1% improvement in SRCC for overall quality prediction and generating explanations highly aligned with human judgments.