Score
Designs and produces per-pixel image quality annotations, including annotation protocols, labeling tools, and benchmark datasets that record pixel-wise visibility or quality values. Builds and analyzes pixel-wise quality metrics and evaluation pipelines (e.g., pixel-level FIQA and localized degradation assessment) to standardize and compare task-agnostic, localized image-quality measurements.
Image quality assessment (IQA) faces core challenges including poor scene adaptability, limited interpretability, and difficulties in engineering deployment. This paper presents a systematic survey of recent IQA advances, categorizing methods—classical metrics (e.g., PSNR, SSIM), machine learning approaches (e.g., SVM, RF), and deep models (e.g., CNN, ViT)—by application scenario. Crucially, it is the first to integrate distortion-specific requirements with practical constraints—including utility, interpretability, and implementation simplicity—into a unified methodological framework. We propose an application-oriented IQA evaluation taxonomy and construct a comprehensive landscape spanning general-purpose and domain-specific methods. Key technical bottlenecks are explicitly identified, and empirically grounded future research directions are provided. The work establishes a benchmark reference for the IQA community, balancing theoretical rigor with actionable engineering guidance.
Current fundus image quality assessment methods predominantly rely on image-level labels, which struggle to quantify localized degradations and often lack task-agnosticism and interpretability. To address these limitations, this work proposes a pixel-level quality assessment approach grounded in the visibility of anatomical structures. The authors introduce FunPiQ, the first pixel-level annotated benchmark for fundus image quality, and develop EFIQA-CP, an inherently interpretable-by-design CNN model. Leveraging non-negative positive-unlabeled learning to generate high-quality pseudo-labels for training, the proposed method significantly outperforms both post-hoc explanation of classification models and anomaly detection baselines across multiple experiments, thereby demonstrating the efficacy and superiority of pixel-level quality evaluation.
Traditional image quality assessment relies heavily on subjective mean opinion scores, which incur high annotation costs and lack local interpretability. To address these limitations, this work proposes a label-free, relational, and directional image quality assessment method. It leverages a self-supervised synthetic distortion engine to generate training data and integrates a spatially aware, disentangled distortion map prediction mechanism with a contrastive learning–based relational scoring network. The proposed approach accurately identifies distortion type, intensity, and orientation, yielding fine-grained and interpretable quality predictions. Furthermore, it enables targeted optimization of image processing algorithms by providing actionable, spatially localized quality feedback.
Image data quality significantly impacts model performance, yet systematic evaluation methodologies remain lacking. This paper proposes an automated image quality assessment pipeline integrating CleanVision and Fastdup, specifically targeting quantitative detection of degradation artifacts—including blur and excessive scaling. We introduce a novel automatic threshold selection mechanism that robustly identifies images with compromised critical visual features without manual parameter tuning, and enhance near-duplicate sample deduplication. Under a binary classification evaluation framework, our method achieves an F1-score of 0.9468 (+0.2674) for single-distortion detection and 0.8557 (+0.1110) for dual-distortion detection; near-duplicate detection F1 improves to 0.7928 (+0.3352). These results demonstrate substantial gains in both accuracy and generalizability of data cleaning.
This study addresses the failure of objective image quality assessment (IQA) metrics near the just-noticeable difference (JND) threshold in high-fidelity image compression, where subtle compression artifacts evade reliable detection. We systematically evaluate the sensitivity and reliability of mainstream IQA metrics to such fine-grained distortions. To this end, we propose Z-RMSE—a metric incorporating subjective rating uncertainty—and design a novel statistical evaluation framework grounded in hypothesis testing. Furthermore, we construct and publicly release the first benchmark dataset dedicated to high-fidelity compression, comprising the full-range JPEG AIC-3 dataset, a JND-subset, cropping-effect analysis, and integrated evaluation tools. Experiments reveal that existing metrics suffer from overfitting and insufficient discriminability below the JND threshold. Our approach significantly improves consistency and robustness in fine-grained distortion assessment, providing a reproducible benchmark, a principled statistical framework, and open-source infrastructure for next-generation IQA research.
Existing VLM-based image quality assessment (IQA) methods suffer from poor generalization and are hindered by insufficient data scale and quality, limiting their applicability in real-world scenarios. To address these limitations, we propose the first unified IQA framework supporting multi-task (distortion identification, instantaneous scoring, quality attribution reasoning), multi-granularity (concise vs. detailed descriptions), and both full-reference and no-reference settings. We introduce DQ-495K, a large-scale, high-quality descriptive IQA dataset, featuring three novel techniques: ground-truth-guided synthetic data generation, native-resolution preservation, and response confidence filtering with calibration. Our framework adopts an end-to-end VLM training paradigm. Extensive experiments demonstrate state-of-the-art performance across all three tasks—outperforming conventional score-based methods, existing VLM-IQA models, and GPT-4V. Further validation on web image assessment and generative output ranking confirms its strong cross-domain generalization capability.
This work addresses the challenge that existing no-reference image quality assessment (IQA) methods struggle to accurately evaluate low-level artifacts introduced by camera image signal processors (ISPs), while full-reference metrics require pristine reference images that are often unavailable. To overcome this limitation, we propose a novel framework that leverages a single sRGB image along with its ISO metadata to synthesize a proxy reference image, enabling the computation of standard full-reference metrics such as PSNR, SSIM, and LPIPS without access to a ground-truth reference. By combining synthetic data pretraining with lightweight LoRA fine-tuning, our method rapidly adapts to diverse ISP configurations and significantly outperforms conventional no-reference IQA approaches and direct regression strategies on real-world camera data, achieving notable improvements in both metric estimation accuracy and ranking consistency.
This work addresses the limitations of existing image quality assessment (IQA) methods, which rely on static, single-pass scoring and fail to capture the dynamic, localized nature of human visual inspection. To overcome this, the authors propose a tool-augmented active IQA framework that reformulates the assessment process into three stages: structured observation, tool-assisted scrutiny, and calibrated scoring. For the first time, interactive viewing tools—namely a magnifier and a gamma corrector—are introduced to enhance the visual language model’s sensitivity to local artifacts and fine details. A batch-aware training strategy is further devised to improve tool invocation efficiency. The proposed method achieves state-of-the-art performance across multiple IQA benchmarks, attaining a PLCC of 0.854 on the CLIVE dataset.
Existing image quality assessment methods overlook the core requirements of scientific images—namely, scientific correctness and logical completeness—and instead focus narrowly on perceptual fidelity or text-image alignment. This work proposes SIQA, the first multidimensional quality assessment framework tailored for scientific images, which systematically decomposes quality into a knowledge dimension (scientific validity and completeness) and a perceptual dimension (cognitive clarity and disciplinary normativity). Two evaluation protocols are introduced: a multiple-choice semantic understanding task (SIQA-U) and an expert-score alignment task (SIQA-S). Leveraging multimodal large language models and an expert-annotated benchmark, experiments reveal that current models achieve strong score alignment (SIQA-S) but exhibit significant deficiencies in semantic understanding (SIQA-U), with fine-tuning yielding only marginal gains in comprehension. These findings underscore the necessity of multidimensional evaluation and validate the effectiveness of SIQA.
This work addresses the unclear commercial value of current image generation models in real-world design scenarios by proposing the first framework that directly links image quality assessment to human payment decisions. The authors introduce ServImageBench, a dataset comprising 1.07k commercial tasks and 2.05k deliverables, along with ServImageScore—a multidimensional evaluation metric encompassing functional requirements, visual quality, and business necessity. Leveraging 33k human-annotated images, they train ServImageModel, a payment prediction model that achieves 82.00% accuracy in forecasting whether users are willing to pay for a given image. The model further outputs calibrated payment probabilities, offering an effective quantitative measure of an image’s commercial viability.
Existing no-reference image quality assessment methods suffer from critical limitations in multi-resolution scenarios, including loss of essential quality cues, poor cross-resolution generalization, difficulty in jointly training on heterogeneous data, and high computational overhead. This work proposes ReLIQS, a novel model that achieves resolution-agnostic quality prediction for the first time by integrating multi-scale patch sampling, a CLIP vision backbone, a perceptual importance estimator, and a latent quality-axis aggregation module. ReLIQS preserves original-resolution quality signals while enabling robust cross-resolution generalization and joint training across heterogeneous MOS scales, further enhanced by a quality-aware saliency mechanism that dynamically selects informative regions. Experiments demonstrate that ReLIQS consistently outperforms CNN-, CLIP-, and MLLM-based baselines on diverse benchmarks encompassing real-world, synthetic, and AIGC-generated images, achieving superior performance at comparable or lower computational cost.