Score
Designs and builds no-reference (blind) image quality assessment systems that jointly predict overall perceptual quality (e.g., MOS) and multiple attribute- or dimension-specific quality scores, producing quantitative image quality metrics and rankings without access to reference images. Implements shared quality-aware visual encoders (often leveraging pre‑trained features) with separate regression or prediction heads per task, trains with multi-task losses that exploit correlated objectives (e.g., PLCC-based losses), and can be target-aware by predicting global, target, and background quality or detecting enhancement artifacts.
Image quality assessment (IQA) faces core challenges including poor scene adaptability, limited interpretability, and difficulties in engineering deployment. This paper presents a systematic survey of recent IQA advances, categorizing methods—classical metrics (e.g., PSNR, SSIM), machine learning approaches (e.g., SVM, RF), and deep models (e.g., CNN, ViT)—by application scenario. Crucially, it is the first to integrate distortion-specific requirements with practical constraints—including utility, interpretability, and implementation simplicity—into a unified methodological framework. We propose an application-oriented IQA evaluation taxonomy and construct a comprehensive landscape spanning general-purpose and domain-specific methods. Key technical bottlenecks are explicitly identified, and empirically grounded future research directions are provided. The work establishes a benchmark reference for the IQA community, balancing theoretical rigor with actionable engineering guidance.
Traditional image quality assessment relies heavily on subjective mean opinion scores, which incur high annotation costs and lack local interpretability. To address these limitations, this work proposes a label-free, relational, and directional image quality assessment method. It leverages a self-supervised synthetic distortion engine to generate training data and integrates a spatially aware, disentangled distortion map prediction mechanism with a contrastive learning–based relational scoring network. The proposed approach accurately identifies distortion type, intensity, and orientation, yielding fine-grained and interpretable quality predictions. Furthermore, it enables targeted optimization of image processing algorithms by providing actionable, spatially localized quality feedback.
To address the limited generalizability and robustness of no-reference image quality assessment (NR-IQA) models—stemming from strong subjectivity in human perception and the complexity of real-world distortions—this paper proposes a quality-aware pretraining and meta-learning synergistic framework. The framework uniquely integrates quality-oriented self-supervised pretraining, a customized quality-aware loss function, and a meta-learning-driven multi-model ensemble mechanism. Built upon a CNN-based feature extractor, it is jointly trained and validated on multiple benchmark datasets: LIVECD, KonIQ-10K, and BIQ2021. Experimental results demonstrate state-of-the-art performance: on in-distribution datasets, it achieves PLCC scores of 0.9885/0.9702/0.884 and SROCC scores of 0.9812/0.9658/0.8765; in cross-dataset evaluation, it attains PLCC/SROCC ranging from 0.6721–0.8023 / 0.6515–0.7805—significantly outperforming existing methods. The framework substantially enhances model robustness to authentic distortions and cross-domain generalization capability.
Existing no-reference image quality assessment (NR-IQA) methods often neglect local manifold structures, leading to insufficient discriminability for challenging distortions. To address this, we propose a contrastive learning framework that explicitly preserves local manifold geometry. First, we introduce non-salient regions from the same image as intra-image negative samples to enhance local discriminability. Second, we design a saliency-guided dual-branch mutual learning mechanism to adaptively emphasize critical visual regions. Third, we integrate multi-scale cropping sampling with a local manifold-constrained contrastive loss. Extensive experiments on seven benchmark datasets demonstrate state-of-the-art performance: PLCC scores of 0.942 on TID2013 and 0.914 on LIVEC—surpassing all prior methods. Crucially, our approach significantly improves perceptual modeling of structurally distorted and noisy images, validating its effectiveness for difficult distortion cases.
This work addresses the long-standing disconnection between image quality assessment (IQA) and exemplar-guided image processing. We propose DisQUE, a self-supervised disentangled representation learning framework that unifies both tasks within a shared content-appearance feature space for the first time. DisQUE employs a dual-stream network architecture coupled with self-supervised contrastive learning to achieve unsupervised disentanglement of content and appearance representations. It introduces an appearance-transfer-based quality prediction module and an exemplar-driven feature modulation mechanism to support appearance editing. On standard IQA benchmarks (e.g., LIVE, TID2013), DisQUE achieves state-of-the-art zero-shot cross-distortion quality prediction. In HDR tone mapping, it faithfully reproduces target appearances from only a few exemplars, demonstrating strong generalization. Our core contribution lies in establishing a unified disentangled representation that jointly supports both IQA and exemplar-guided processing—overcoming the limitations of conventional single-task modeling paradigms.
Existing no-reference image quality assessment (NR-IQA) methods rely on semantic backbone networks for feature extraction; however, their outputs often contain semantically irrelevant—or even detrimental—noise, leading to substantial quality score discrepancies between image pairs with small feature distances. To address this, we propose Quality-aware Feature Matching IQM (QFM-IQM), the first NR-IQA framework to introduce adversarial semantic noise matching: it identifies such noise by constructing image pairs that are quality-similar but semantically dissimilar. QFM-IQM explicitly models feature sensitivity to semantic noise and adaptively suppresses redundant or harmful channels via channel-wise gating. Additionally, knowledge distillation is integrated to enhance generalization across diverse distortion types. Evaluated on eight standard IQA benchmarks, QFM-IQM consistently outperforms state-of-the-art methods, achieving significant improvements in prediction accuracy, cross-dataset generalization, and robustness to semantic confounders.
This work addresses the challenges in no-reference image quality assessment arising from the difficulty of effectively fusing natural scene statistics (NSS) features with visual language model (VLM) embeddings, and the fact that the contribution of each feature varies dynamically across distortion types. To this end, the authors propose a distortion-aware dynamic fusion framework that adaptively weights 138-dimensional NSS features alongside SigLIP and CLIP-H embeddings via a multiplicative gating mechanism conditioned on input image content, without requiring fine-tuning of the VLM backbone. This approach achieves the first content-conditioned dynamic fusion of NSS and VLM features, with gating weights aligning closely with human distortion analyses. The method sets new state-of-the-art results on KonIQ-10k, KADID-10k, and LIVE Challenge, achieving an SROCC of 0.9715 and PLCC of 0.9733 on KADID-10k, and demonstrates that NSS features are most influential for noise and color distortion types.
This work addresses the challenge that existing no-reference image quality assessment (IQA) methods struggle to accurately evaluate low-level artifacts introduced by camera image signal processors (ISPs), while full-reference metrics require pristine reference images that are often unavailable. To overcome this limitation, we propose a novel framework that leverages a single sRGB image along with its ISO metadata to synthesize a proxy reference image, enabling the computation of standard full-reference metrics such as PSNR, SSIM, and LPIPS without access to a ground-truth reference. By combining synthetic data pretraining with lightweight LoRA fine-tuning, our method rapidly adapts to diverse ISP configurations and significantly outperforms conventional no-reference IQA approaches and direct regression strategies on real-world camera data, achieving notable improvements in both metric estimation accuracy and ranking consistency.
This work addresses the limitations of existing image quality assessment (IQA) methods, which rely on static, single-pass scoring and fail to capture the dynamic, localized nature of human visual inspection. To overcome this, the authors propose a tool-augmented active IQA framework that reformulates the assessment process into three stages: structured observation, tool-assisted scrutiny, and calibrated scoring. For the first time, interactive viewing tools—namely a magnifier and a gamma corrector—are introduced to enhance the visual language model’s sensitivity to local artifacts and fine details. A batch-aware training strategy is further devised to improve tool invocation efficiency. The proposed method achieves state-of-the-art performance across multiple IQA benchmarks, attaining a PLCC of 0.854 on the CLIVE dataset.
Existing no-reference image quality assessment methods suffer from critical limitations in multi-resolution scenarios, including loss of essential quality cues, poor cross-resolution generalization, difficulty in jointly training on heterogeneous data, and high computational overhead. This work proposes ReLIQS, a novel model that achieves resolution-agnostic quality prediction for the first time by integrating multi-scale patch sampling, a CLIP vision backbone, a perceptual importance estimator, and a latent quality-axis aggregation module. ReLIQS preserves original-resolution quality signals while enabling robust cross-resolution generalization and joint training across heterogeneous MOS scales, further enhanced by a quality-aware saliency mechanism that dynamically selects informative regions. Experiments demonstrate that ReLIQS consistently outperforms CNN-, CLIP-, and MLLM-based baselines on diverse benchmarks encompassing real-world, synthetic, and AIGC-generated images, achieving superior performance at comparable or lower computational cost.
This work addresses the challenges in blind ultra-high-definition (UHD) image quality assessment, where full-resolution inference is computationally prohibitive and naive downsampling or isolated patch cropping fails to capture scale-sensitive distortions and global-local dependencies. To overcome these limitations, the authors propose the first graph neural network–based approach that encodes image patches as nodes and constructs a hybrid k-nearest neighbor graph based on spatial proximity and feature similarity. Contextual information is propagated via residual graph convolutions, and regional evidence is aggregated through gated attention pooling to predict overall image quality. Departing from conventional independent-processing assumptions, the method explicitly models inter-regional structural dependencies and introduces a multi-objective loss function with exponential moving average normalization to jointly optimize regression accuracy, correlation, and ranking performance. On the UHD-IQA benchmark, it achieves state-of-the-art results with PLCC = 0.7784, SRCC = 0.8019, and RMSE = 0.0519—the lowest RMSE reported to date—demonstrating significantly improved absolute quality prediction accuracy.