perceptual loss design

Design and implement perceptual loss functions and associated optimization/tuning procedures that quantify and penalize perceptual discrepancies in visual signals using feature‑space similarity metrics, region weighting, color‑angle terms, luminance‑edge regularization, and related color/edge penalties. Integrate these losses as reconstruction or regularization components to preserve fine structural topology, enforce hue/brightness consistency, reduce color bias, and suppress local visual artifacts.

perceptuallossdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Perceptually Inspired Variational Framework for Color Enhancement

Nov 28, 2025
RP
Rodrigo Palma-Amestoy
🏛️ Universidad de Chile | Università di Milano | Universitat Pompeu Fabra

This paper addresses the longstanding challenge in image color enhancement of simultaneously achieving visual naturalness and computational efficiency. We propose a human perception-inspired variational framework. Methodologically, we first systematically formulate three fundamental principles for perception-inspired energy functionals, then construct three theoretically grounded explicit functionals—respectively modeling color contrast, chromatic distribution dispersion, and perceptual consistency. The optimization is performed via gradient descent, augmented by a generic acceleration strategy based on the fast Fourier transform (FFT), reducing computational complexity from $O(N^2)$ to $O(N log N)$. Experiments demonstrate that our method significantly outperforms conventional approaches across diverse images: it enhances color contrast and distribution rationality while better preserving fine details and visual naturalness, all at substantially reduced computational cost.

Characterizing model behavior for image contrast and dispersionDeveloping perceptually inspired variational color enhancement algorithmsReducing computational cost from O(N^2) to O(N log N)

VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression

Dec 17, 2025
KS
Kyle Sargent
🏛️ Stanford University | Google Research | Google DeepMind

Traditional image compression evaluation relies on distortion metrics such as MSE, which exhibit significant misalignment with human perceptual judgment. To address this, we propose VLIC—the first perception-aligned compression framework leveraging diffusion models and frozen vision-language models (VLMs, e.g., CLIP or LLaVA) as zero-shot preference discriminators. Crucially, VLIC bypasses fine-tuning or distillation; instead, it employs VLMs to perform binary Alternative Forced Choice (AFC) comparisons on compressed image pairs, generating preference-based reward signals to guide diffusion model post-training. This paradigm establishes the first native, parameter-free VLM-driven perceptual guidance for compression. Extensive experiments demonstrate that VLIC achieves state-of-the-art performance across perceptual metrics—including LPIPS and DISTS—as well as in large-scale user studies, significantly outperforming both classical and learned compression methods.

Evaluates image compression alignment with human perceptionProposes diffusion-based compression trained with VLM binary preferencesUses vision-language models for perceptual judgments zero-shot

From Images to Perception: Emergence of Perceptual Properties by Reconstructing Images

Aug 14, 2025
PH
Pablo Hernández-Cámara
🏛️ University of Valencia

This study investigates whether human-like visual perception can emerge unsupervised from natural image statistics. Method: We propose PerceptNet, a biologically inspired architecture modeling retinal–V1 processing, trained end-to-end via joint optimization of multiple self-supervised objectives: image reconstruction (autoencoding), denoising, deblurring, and sparse regularization—without any perceptual labels. Contribution/Results: The learned encoding-layer representations achieve remarkably high alignment with human subjective perceptual judgments (Pearson’s *r* > 0.9), matching human performance under moderate noise, blur, and sparsity constraints. Critically, this correspondence emerges purely from statistical regularities in natural images, without task-specific supervision. Our work provides the first systematic computational evidence that biologically grounded models can spontaneously develop human-aligned perceptual metrics solely through unsupervised learning on image statistics. These findings substantiate efficient coding theories of early vision and establish a new paradigm for modeling perceptual representation emergence.

Develops bio-inspired model for image reconstruction tasksExplores how human visual perception emerges from image statisticsShows model aligns with human perception without supervision

A prevalent yet long-overlooked issue in supervised low-light image enhancement (LLIE) is brightness mismatch between enhanced outputs and ground-truth images, leading to training bias. This paper is the first to systematically identify and analyze this phenomenon. We propose GT-mean loss—a principled extension of standard L1/L2 losses—by probabilistically modeling the distribution of image mean intensities and enforcing explicit ground-truth mean constraints. The loss is plug-and-play: it integrates seamlessly into existing supervised frameworks without introducing additional parameters or measurable computational overhead, and remains compatible with mainstream LLIE methods. Extensive experiments across multiple benchmarks and base models demonstrate that GT-mean consistently improves PSNR and SSIM while effectively mitigating brightness mismatch. Our approach provides a simple, parameter-free, and broadly applicable brightness calibration solution for supervised LLIE.

Addresses brightness mismatch in low-light image enhancementEnhances performance across multiple methods and datasetsProposes GT-mean loss to improve model training accuracy

Training Neural Networks on RAW and HDR Images for Restoration Tasks

Dec 06, 2023
LL
Lei Luo
🏛️ Meta | University of Cambridge

This work addresses the limited effectiveness of linear color space modeling in RAW/HDR image restoration (denoising, deblurring, super-resolution). We systematically validate and propose a novel training paradigm that replaces conventional linear color spaces with display-encoded gamuts—specifically PQ, PU21, and mu-law—to better align neural network optimization with human visual perception. Our key contribution is the first empirical demonstration that perceptually uniform display encodings significantly accelerate model convergence and improve reconstruction quality. By integrating standard CNN architectures with perceptually calibrated loss functions, we achieve end-to-end optimization. Across denoising, deblurring, and super-resolution tasks, our approach yields consistent PSNR gains of 2–9 dB. Crucially, it preserves physical interpretability of RAW/HDR data while better satisfying both perceptual fidelity and deep learning optimization requirements. This establishes a generalizable, perception-aware training framework for high-dynamic-range image restoration.

Comparing display-encoded vs. linear color spaces for training.Evaluating perceptual uniformity impact on denoising, deblurring, super-resolution.Optimizing neural network training for RAW and HDR image restoration.

Latest Papers

What's happening recently
View more

This study addresses the misalignment between existing color difference metrics and human perception by training regression models on human similarity judgment data to predict perceived color differences. We propose COLIBRI features, which integrate numerical coordinates with fuzzy linguistic categories, and evaluate multiple machine learning algorithms. Our findings indicate that color representation features are more critical than algorithm selection. Specifically, a LightGBM model incorporating the proposed feature representation achieves an R² of 0.703, significantly outperforming conventional RGB and HSI methods. This approach effectively enhances both the accuracy and perceptual consistency of color difference assessment.

color difference metricshuman perceptionhuman similarity judgments

This study addresses the generalization bias in post-training quantization of large vision-language models (LVLMs) caused by an over-reliance on reconstruction loss. To overcome this limitation, we propose Balanced Fitting, a framework that departs from the conventional error-minimization paradigm by exploiting the regularization benefits that quantization confers upon specific layers and modalities. Through fine-grained evaluation of component-wise quantization effects, a hybrid fitting strategy, and joint weight-activation quantization, our approach dynamically balances accuracy preservation with regularization gains. Extensive experiments demonstrate that the proposed method significantly outperforms existing baselines across diverse LVLM architectures. These findings compellingly establish that low reconstruction loss does not necessarily translate to superior downstream performance, thereby introducing a new paradigm for multimodal model quantization.

GeneralizationLarge Vision-Language ModelsPost-Training Quantization

本文综述了率-失真-感知框架,从数学理论到实际应用的演变,探讨了Blau-Michaeli函数及其计算方法,为下一代感知压缩系统提供严谨基础。

Blau--Michaeli functionperceptual compression systemsrate-distortion-perception

Traditional research on graphical perception has predominantly evaluated visualizations from an encoding perspective, often overlooking the fact that the human visual system processes pixel-based images, thereby creating a disconnect between evaluation and actual perception. This work proposes treating visualizations as images and, for the first time, systematically integrates summary statistic vision theory by employing computational vision models that take pixels as input to model the perceptual process from the decoding end. The approach not only successfully reproduces established findings in graphical perception but also sensitively predicts perceptual changes induced by subtle variations in data distributions or design choices, demonstrating the effectiveness and potential of image-based vision models for evaluating visualizations.

graphical perceptionhuman visionimage-based modeling

为解决像素空间流模型训练中低频信号主导优化的问题,提出了一种频谱平衡目标函数f-loss,并结合频率和像素监督来加速收敛并提高生成质量。

optimization signalperceptually important structurespixel-space flow models

Hot Scholars

RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression
LS

Li Song

Professor of Electronic Engineering, Shanghai Jiao Tong University
Video CodingImage ProcessingComputer Vision
WZ

Wangmeng Zuo

School of Computer Science and Technology, Harbin Institute of Technology
Computer VisionImage ProcessingGenerative AIDeep Learning
MV

Marcos V. Conde

Ph.D. Researcher, University of Würzburg, Sony PlayStation, CIDAUT
Artificial IntelligenceDeep LearningComputer VisionImage Processing
ZJ

Zhaoyang Jia

University of Science and Technology of China
Video compressiondigital watermarking