Score
Design and implement perceptual loss functions and associated optimization/tuning procedures that quantify and penalize perceptual discrepancies in visual signals using feature‑space similarity metrics, region weighting, color‑angle terms, luminance‑edge regularization, and related color/edge penalties. Integrate these losses as reconstruction or regularization components to preserve fine structural topology, enforce hue/brightness consistency, reduce color bias, and suppress local visual artifacts.
This paper addresses the lack of systematic guidance for loss function selection and design in deep learning. We propose the first cross-task loss taxonomy—covering vision, time-series, and tabular data—and unifying discriminative and generative paradigms. Through comprehensive survey analysis, rigorous mathematical modeling, and multi-scenario empirical evaluation, we characterize the applicability boundaries and failure modes of 12 mainstream losses, identifying three fundamental challenges: low computational efficiency, gradient instability, and poor adaptability to real-world constraints. Building on these insights, we formulate next-generation loss design principles centered on robustness, interpretability, and adaptivity. Furthermore, we deliver a practical, industry-deployment-oriented loss selection guide—grounded in empirical evidence and operational feasibility—to bridge the gap between theoretical design and practical application.
This paper addresses the longstanding challenge in image color enhancement of simultaneously achieving visual naturalness and computational efficiency. We propose a human perception-inspired variational framework. Methodologically, we first systematically formulate three fundamental principles for perception-inspired energy functionals, then construct three theoretically grounded explicit functionals—respectively modeling color contrast, chromatic distribution dispersion, and perceptual consistency. The optimization is performed via gradient descent, augmented by a generic acceleration strategy based on the fast Fourier transform (FFT), reducing computational complexity from $O(N^2)$ to $O(N log N)$. Experiments demonstrate that our method significantly outperforms conventional approaches across diverse images: it enhances color contrast and distribution rationality while better preserving fine details and visual naturalness, all at substantially reduced computational cost.
Traditional image compression evaluation relies on distortion metrics such as MSE, which exhibit significant misalignment with human perceptual judgment. To address this, we propose VLIC—the first perception-aligned compression framework leveraging diffusion models and frozen vision-language models (VLMs, e.g., CLIP or LLaVA) as zero-shot preference discriminators. Crucially, VLIC bypasses fine-tuning or distillation; instead, it employs VLMs to perform binary Alternative Forced Choice (AFC) comparisons on compressed image pairs, generating preference-based reward signals to guide diffusion model post-training. This paradigm establishes the first native, parameter-free VLM-driven perceptual guidance for compression. Extensive experiments demonstrate that VLIC achieves state-of-the-art performance across perceptual metrics—including LPIPS and DISTS—as well as in large-scale user studies, significantly outperforming both classical and learned compression methods.
This study investigates whether human-like visual perception can emerge unsupervised from natural image statistics. Method: We propose PerceptNet, a biologically inspired architecture modeling retinal–V1 processing, trained end-to-end via joint optimization of multiple self-supervised objectives: image reconstruction (autoencoding), denoising, deblurring, and sparse regularization—without any perceptual labels. Contribution/Results: The learned encoding-layer representations achieve remarkably high alignment with human subjective perceptual judgments (Pearson’s *r* > 0.9), matching human performance under moderate noise, blur, and sparsity constraints. Critically, this correspondence emerges purely from statistical regularities in natural images, without task-specific supervision. Our work provides the first systematic computational evidence that biologically grounded models can spontaneously develop human-aligned perceptual metrics solely through unsupervised learning on image statistics. These findings substantiate efficient coding theories of early vision and establish a new paradigm for modeling perceptual representation emergence.
A prevalent yet long-overlooked issue in supervised low-light image enhancement (LLIE) is brightness mismatch between enhanced outputs and ground-truth images, leading to training bias. This paper is the first to systematically identify and analyze this phenomenon. We propose GT-mean loss—a principled extension of standard L1/L2 losses—by probabilistically modeling the distribution of image mean intensities and enforcing explicit ground-truth mean constraints. The loss is plug-and-play: it integrates seamlessly into existing supervised frameworks without introducing additional parameters or measurable computational overhead, and remains compatible with mainstream LLIE methods. Extensive experiments across multiple benchmarks and base models demonstrate that GT-mean consistently improves PSNR and SSIM while effectively mitigating brightness mismatch. Our approach provides a simple, parameter-free, and broadly applicable brightness calibration solution for supervised LLIE.
This work addresses the limited effectiveness of linear color space modeling in RAW/HDR image restoration (denoising, deblurring, super-resolution). We systematically validate and propose a novel training paradigm that replaces conventional linear color spaces with display-encoded gamuts—specifically PQ, PU21, and mu-law—to better align neural network optimization with human visual perception. Our key contribution is the first empirical demonstration that perceptually uniform display encodings significantly accelerate model convergence and improve reconstruction quality. By integrating standard CNN architectures with perceptually calibrated loss functions, we achieve end-to-end optimization. Across denoising, deblurring, and super-resolution tasks, our approach yields consistent PSNR gains of 2–9 dB. Crucially, it preserves physical interpretability of RAW/HDR data while better satisfying both perceptual fidelity and deep learning optimization requirements. This establishes a generalizable, perception-aware training framework for high-dynamic-range image restoration.
This study addresses the misalignment between existing color difference metrics and human perception by training regression models on human similarity judgment data to predict perceived color differences. We propose COLIBRI features, which integrate numerical coordinates with fuzzy linguistic categories, and evaluate multiple machine learning algorithms. Our findings indicate that color representation features are more critical than algorithm selection. Specifically, a LightGBM model incorporating the proposed feature representation achieves an R² of 0.703, significantly outperforming conventional RGB and HSI methods. This approach effectively enhances both the accuracy and perceptual consistency of color difference assessment.
This study addresses the generalization bias in post-training quantization of large vision-language models (LVLMs) caused by an over-reliance on reconstruction loss. To overcome this limitation, we propose Balanced Fitting, a framework that departs from the conventional error-minimization paradigm by exploiting the regularization benefits that quantization confers upon specific layers and modalities. Through fine-grained evaluation of component-wise quantization effects, a hybrid fitting strategy, and joint weight-activation quantization, our approach dynamically balances accuracy preservation with regularization gains. Extensive experiments demonstrate that the proposed method significantly outperforms existing baselines across diverse LVLM architectures. These findings compellingly establish that low reconstruction loss does not necessarily translate to superior downstream performance, thereby introducing a new paradigm for multimodal model quantization.
本文综述了率-失真-感知框架,从数学理论到实际应用的演变,探讨了Blau-Michaeli函数及其计算方法,为下一代感知压缩系统提供严谨基础。
Traditional research on graphical perception has predominantly evaluated visualizations from an encoding perspective, often overlooking the fact that the human visual system processes pixel-based images, thereby creating a disconnect between evaluation and actual perception. This work proposes treating visualizations as images and, for the first time, systematically integrates summary statistic vision theory by employing computational vision models that take pixels as input to model the perceptual process from the decoding end. The approach not only successfully reproduces established findings in graphical perception but also sensitively predicts perceptual changes induced by subtle variations in data distributions or design choices, demonstrating the effectiveness and potential of image-based vision models for evaluating visualizations.
为解决像素空间流模型训练中低频信号主导优化的问题,提出了一种频谱平衡目标函数f-loss,并结合频率和像素监督来加速收敛并提高生成质量。