multi-task image quality assessment

Designs and builds no-reference (blind) image quality assessment systems that jointly predict overall perceptual quality (e.g., MOS) and multiple attribute- or dimension-specific quality scores, producing quantitative image quality metrics and rankings without access to reference images. Implements shared quality-aware visual encoders (often leveraging pre‑trained features) with separate regression or prediction heads per task, trains with multi-task losses that exploit correlated objectives (e.g., PLCC-based losses), and can be target-aware by predicting global, target, and background quality or detecting enhancement artifacts.

multi-taskimagequalityassessment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Traditional image quality assessment relies heavily on subjective mean opinion scores, which incur high annotation costs and lack local interpretability. To address these limitations, this work proposes a label-free, relational, and directional image quality assessment method. It leverages a self-supervised synthetic distortion engine to generate training data and integrates a spatially aware, disentangled distortion map prediction mechanism with a contrastive learning–based relational scoring network. The proposed approach accurately identifies distortion type, intensity, and orientation, yielding fine-grained and interpretable quality predictions. Furthermore, it enables targeted optimization of image processing algorithms by providing actionable, spatially localized quality feedback.

image quality assessmentinterpretable feedbacklocalized distortion analysis

MetaQAP -- A Meta-Learning Approach for Quality-Aware Pretraining in Image Quality Assessment

Jun 19, 2025
MA
Muhammad Azeem Aslam
🏛️ Xi’an Eurasia University | Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences | Xidian University | University of Engineering and Technology Lahore | University of Central Punjab | Xi’an Jiaotong University

To address the limited generalizability and robustness of no-reference image quality assessment (NR-IQA) models—stemming from strong subjectivity in human perception and the complexity of real-world distortions—this paper proposes a quality-aware pretraining and meta-learning synergistic framework. The framework uniquely integrates quality-oriented self-supervised pretraining, a customized quality-aware loss function, and a meta-learning-driven multi-model ensemble mechanism. Built upon a CNN-based feature extractor, it is jointly trained and validated on multiple benchmark datasets: LIVECD, KonIQ-10K, and BIQ2021. Experimental results demonstrate state-of-the-art performance: on in-distribution datasets, it achieves PLCC scores of 0.9885/0.9702/0.884 and SROCC scores of 0.9812/0.9658/0.8765; in cross-dataset evaluation, it attains PLCC/SROCC ranging from 0.6721–0.8023 / 0.6515–0.7805—significantly outperforming existing methods. The framework substantially enhances model robustness to authentic distortions and cross-domain generalization capability.

Addressing subjective human perception in image quality assessmentEnhancing generalizability of no-reference IQA modelsOvercoming complexity of real-world image distortions in IQA

Local Manifold Learning for No-Reference Image Quality Assessment

Jun 27, 2024
TG
Timin Gao
🏛️ Xiamen University | Harbin Institute of Technology | Tencent Youtu Lab

Existing no-reference image quality assessment (NR-IQA) methods often neglect local manifold structures, leading to insufficient discriminability for challenging distortions. To address this, we propose a contrastive learning framework that explicitly preserves local manifold geometry. First, we introduce non-salient regions from the same image as intra-image negative samples to enhance local discriminability. Second, we design a saliency-guided dual-branch mutual learning mechanism to adaptively emphasize critical visual regions. Third, we integrate multi-scale cropping sampling with a local manifold-constrained contrastive loss. Extensive experiments on seven benchmark datasets demonstrate state-of-the-art performance: PLCC scores of 0.942 on TID2013 and 0.914 on LIVEC—surpassing all prior methods. Crucially, our approach significantly improves perceptual modeling of structurally distorted and noisy images, validating its effectiveness for difficult distortion cases.

Enhancing discriminative capability through contrastive local manifold learningImproving recognition of visually salient regions via mutual learningOvercoming neglect of local manifold structures in image quality assessment

This work addresses the long-standing disconnection between image quality assessment (IQA) and exemplar-guided image processing. We propose DisQUE, a self-supervised disentangled representation learning framework that unifies both tasks within a shared content-appearance feature space for the first time. DisQUE employs a dual-stream network architecture coupled with self-supervised contrastive learning to achieve unsupervised disentanglement of content and appearance representations. It introduces an appearance-transfer-based quality prediction module and an exemplar-driven feature modulation mechanism to support appearance editing. On standard IQA benchmarks (e.g., LIVE, TID2013), DisQUE achieves state-of-the-art zero-shot cross-distortion quality prediction. In HDR tone mapping, it faithfully reproduces target appearances from only a few exemplars, demonstrating strong generalization. Our core contribution lies in establishing a unified disentangled representation that jointly supports both IQA and exemplar-guided processing—overcoming the limitations of conventional single-task modeling paradigms.

Create DisQUE model for state-of-the-art image quality assessmentDevelop disentangled representation learning for image appearance and content separationEnable example-guided image processing using learned disentangled features

Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise Sensitivity

Dec 11, 2023
XL
Xudong Li
🏛️ Xiamen University | Beijing Institute of Technology | Tencent Youtu Lab | Ocean University of China

Existing no-reference image quality assessment (NR-IQA) methods rely on semantic backbone networks for feature extraction; however, their outputs often contain semantically irrelevant—or even detrimental—noise, leading to substantial quality score discrepancies between image pairs with small feature distances. To address this, we propose Quality-aware Feature Matching IQM (QFM-IQM), the first NR-IQA framework to introduce adversarial semantic noise matching: it identifies such noise by constructing image pairs that are quality-similar but semantically dissimilar. QFM-IQM explicitly models feature sensitivity to semantic noise and adaptively suppresses redundant or harmful channels via channel-wise gating. Additionally, knowledge distillation is integrated to enhance generalization across diverse distortion types. Evaluated on eight standard IQA benchmarks, QFM-IQM consistently outperforms state-of-the-art methods, achieving significant improvements in prediction accuracy, cross-dataset generalization, and robustness to semantic confounders.

Addressing quality-irrelevant noise in extracted image featuresImproving NR-IQA accuracy via adversarial semantic noise reductionSelecting beneficial features for NR-IQA by removing harmful noise

Latest Papers

What's happening recently
View more

This work addresses the challenges in no-reference image quality assessment arising from the difficulty of effectively fusing natural scene statistics (NSS) features with visual language model (VLM) embeddings, and the fact that the contribution of each feature varies dynamically across distortion types. To this end, the authors propose a distortion-aware dynamic fusion framework that adaptively weights 138-dimensional NSS features alongside SigLIP and CLIP-H embeddings via a multiplicative gating mechanism conditioned on input image content, without requiring fine-tuning of the VLM backbone. This approach achieves the first content-conditioned dynamic fusion of NSS and VLM features, with gating weights aligning closely with human distortion analyses. The method sets new state-of-the-art results on KonIQ-10k, KADID-10k, and LIVE Challenge, achieving an SROCC of 0.9715 and PLCC of 0.9733 on KADID-10k, and demonstrates that NSS features are most influential for noise and color distortion types.

Blind Image Quality AssessmentDistortion-AwareFeature Fusion

This work addresses the challenge that existing no-reference image quality assessment (IQA) methods struggle to accurately evaluate low-level artifacts introduced by camera image signal processors (ISPs), while full-reference metrics require pristine reference images that are often unavailable. To overcome this limitation, we propose a novel framework that leverages a single sRGB image along with its ISO metadata to synthesize a proxy reference image, enabling the computation of standard full-reference metrics such as PSNR, SSIM, and LPIPS without access to a ground-truth reference. By combining synthetic data pretraining with lightweight LoRA fine-tuning, our method rapidly adapts to diverse ISP configurations and significantly outperforms conventional no-reference IQA approaches and direct regression strategies on real-world camera data, achieving notable improvements in both metric estimation accuracy and ranking consistency.

full-reference metricsimage quality assessmentISP pipeline evaluation

This work addresses the limitations of existing image quality assessment (IQA) methods, which rely on static, single-pass scoring and fail to capture the dynamic, localized nature of human visual inspection. To overcome this, the authors propose a tool-augmented active IQA framework that reformulates the assessment process into three stages: structured observation, tool-assisted scrutiny, and calibrated scoring. For the first time, interactive viewing tools—namely a magnifier and a gamma corrector—are introduced to enhance the visual language model’s sensitivity to local artifacts and fine details. A batch-aware training strategy is further devised to improve tool invocation efficiency. The proposed method achieves state-of-the-art performance across multiple IQA benchmarks, attaining a PLCC of 0.854 on the CLIVE dataset.

ArtifactsImage Quality AssessmentLocal Details

Existing no-reference image quality assessment methods suffer from critical limitations in multi-resolution scenarios, including loss of essential quality cues, poor cross-resolution generalization, difficulty in jointly training on heterogeneous data, and high computational overhead. This work proposes ReLIQS, a novel model that achieves resolution-agnostic quality prediction for the first time by integrating multi-scale patch sampling, a CLIP vision backbone, a perceptual importance estimator, and a latent quality-axis aggregation module. ReLIQS preserves original-resolution quality signals while enabling robust cross-resolution generalization and joint training across heterogeneous MOS scales, further enhanced by a quality-aware saliency mechanism that dynamically selects informative regions. Experiments demonstrate that ReLIQS consistently outperforms CNN-, CLIP-, and MLLM-based baselines on diverse benchmarks encompassing real-world, synthetic, and AIGC-generated images, achieving superior performance at comparable or lower computational cost.

Computational EfficiencyHeterogeneous IQA DatasetsNo-reference Image Quality Assessment

This work addresses the challenges in blind ultra-high-definition (UHD) image quality assessment, where full-resolution inference is computationally prohibitive and naive downsampling or isolated patch cropping fails to capture scale-sensitive distortions and global-local dependencies. To overcome these limitations, the authors propose the first graph neural network–based approach that encodes image patches as nodes and constructs a hybrid k-nearest neighbor graph based on spatial proximity and feature similarity. Contextual information is propagated via residual graph convolutions, and regional evidence is aggregated through gated attention pooling to predict overall image quality. Departing from conventional independent-processing assumptions, the method explicitly models inter-regional structural dependencies and introduces a multi-objective loss function with exponential moving average normalization to jointly optimize regression accuracy, correlation, and ranking performance. On the UHD-IQA benchmark, it achieves state-of-the-art results with PLCC = 0.7784, SRCC = 0.8019, and RMSE = 0.0519—the lowest RMSE reported to date—demonstrating significantly improved absolute quality prediction accuracy.

Blind Image Quality AssessmentComputational ComplexityGlobal Scene Context

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression
CH

Chunming He

Duke University | Tsinghua University
Computer VisionMachine LearningBiomedical Image Analysis
CC

Chih-Chung Hsu

Associate Professor of Institute of Intelligent Systems, College of AI, NYCU
Deep learningImage processingcomputer visionimage compression
CM

Chia-Ming Lee

National Yang Ming Chiao Tung University
Computer VisionImage ProcessingInformation ForensicsMultimedia