Score
Designs and evaluates compression codec configurations and encoder parameters for facial images or video to meet bitrate or size constraints while preserving downstream face-recognition accuracy. Builds measurement and comparison pipelines that analyze recognition performance across codecs, bitrate levels, and quality settings and selects optimal encoder settings for target bitrates.
JPEG and JPEG 2000 compression degrades facial image fidelity, leading to reduced face recognition accuracy. Method: This paper proposes an interpretable, pre-processing quality control framework for facial image quality assessment. It unifies detection of artifacts from both compression standards—using PSNR/SSIM-derived weak supervision labels—and trains a lightweight, end-to-end EfficientNetV2 binary classifier without manual annotation. The model supports real-time deployment and achieves a compression-type classification error rate of only 2–3%. Contribution/Results: Extensive evaluation across multiple open-source and commercial face recognition systems demonstrates that filtering out high-artifact images significantly reduces downstream recognition error rates. The proposed method has been integrated into the open-source OFIQ framework, providing actionable, interpretable quality feedback to enhance robustness in face recognition pipelines.
研究探讨了在1kB以下存储人脸图像时如何保留身份信息,通过对比多种编解码器及自定义学习编解码器,解决了低预算下的人脸识别准确性问题。
To address feature degradation and accuracy loss in facial expression recognition (FER) caused by lossy image compression, this paper proposes an end-to-end learnable image compression framework. The method introduces a task-specific joint optimization objective that integrates feature-aware reconstruction loss with classification-guided supervision, enabling adaptive weighting to balance compression fidelity and discriminative feature preservation. Leveraging a deep learning–based compression backbone, the framework supports both standalone fine-tuning and end-to-end joint training. Experiments show that standalone fine-tuning improves FER accuracy by 0.71% while reducing bit-rate by 49.32%; joint optimization further boosts accuracy by 4.04% and reduces bit-rate by 89.12%, maintaining model stability in both compressed and pixel domains. The core contribution is the first integration of FER-driven discriminative constraints into a learnable compression pipeline, achieving high-accuracy recognition under high compression ratios.
Existing video compression methods are primarily optimized for human visual perception and thus often fail to preserve semantic information critical for machine vision tasks. To address this, this paper proposes a machine-vision-oriented neural preprocessing framework. It introduces a learnable preprocessor prior to standard video encoding and pioneers a differentiable virtual codec, enabling end-to-end joint optimization of preprocessing and conventional encoders (e.g., H.264/AVC) without modifying codec standards. A rate–distortion–task loss jointly optimizes bit rate, reconstruction fidelity, and downstream task performance—including object detection and action recognition. Experiments demonstrate that the framework reduces average bit rate by over 15% while maintaining or even improving task accuracy, significantly enhancing semantic fidelity and utility of compressed video for machine vision applications.
Existing video codecs are optimized for human visual perception, neglecting the impact of compression artifacts on machine vision tasks—leading to degraded downstream AI performance. This paper proposes CDRE, a framework that models and compresses “compression-sensitive distortion” in the feature domain as a learnable, transmittable machine-perception embedding. Methodologically, CDRE introduces: (1) the first distortion-driven differentiable embedding paradigm; (2) a co-designed architecture integrating a compression-sensitive feature extractor with a lightweight distortion codec (comprising quantization and entropy modeling); and (3) progressive embedding adaptation and rate–task joint optimization during training. Evaluated on H.264, H.265, and AV1 encoders across object detection and action recognition tasks, CDRE achieves an average mAP gain of 3.2%, with marginal overhead—less than 0.5% bitrate increase and under 1% additional model parameters.
This study addresses the stringent requirement of compressing facial images to under 1024 bytes for storage in temporary travel documents by systematically evaluating multiple image coding formats—including JPEG, JPEG 2000, JPEG XL, JPEG AI, HEIF, AVIF, and WebP—in conjunction with preprocessing strategies such as grayscale conversion, smoothing, and resizing. It presents the first comprehensive comparison of next-generation compression standards under this extreme bitrate constraint, specifically assessing their impact on face recognition performance. For ICAO-compliant and non-compliant images, distinct optimal solutions emerge: JPEG AI achieves the highest recognition accuracy for high-quality images, followed closely by AVIF and WebP. Grayscale conversion improves recognition rates for compliant images, whereas retaining color information proves more beneficial for low-quality inputs.
Existing learned image codecs struggle to simultaneously achieve high perceptual quality and real-time performance on edge devices. This work systematically investigates key modeling choices affecting practicality and proposes a unified optimization framework that integrates differentiable compression, perception-driven loss optimization, and ablation-guided module design. For the first time, it conducts a large-scale neural architecture search (NAS) over millions of backbone configurations under explicit latency and quality constraints. The resulting efficient codec achieves 230 ms encoding and 150 ms decoding for 12MP images on an iPhone 17 Pro Max. In subjective evaluations, it reduces bitrate by 2.3–3× compared to AV1/VVC and further improves upon state-of-the-art learned methods by 20–40% in bitrate savings.
本文针对机器视觉系统的图像压缩问题,通过设定人类可接受的视觉质量上限并优化剩余编码容量以提升机器性能,提出了一种基于约束优化的方法。
Emerging machine vision applications require efficient transmission of neural network intermediate features—rather than pixel data—necessitating a paradigm shift from human-vision-oriented video coding. Method: This paper pioneers the first systematic study of VVC-based feature compression optimized for machine perception, introducing three lightweight coding profiles—Fast, Faster, and Fastest—designed via fine-grained analysis of VVC tool impacts on downstream task accuracy to jointly optimize coding efficiency and inference fidelity. Contribution/Results: Fast achieves a 21.8% reduction in encoding time while improving BD-Rate by 2.96%; Fastest delivers 95.6% encoding acceleration with only a 1.71% BD-Rate degradation. The framework constitutes the first deployable, VVC-based, machine-perception-optimized coding solution proposed to MPEG for standardization under the AI Feature Coding for Machines (FCM) initiative.
This work addresses the challenge of achieving both high fidelity and real-time performance in 3D video conferencing under low-bitrate constraints, where conventional 2D compression discards critical geometric details and implicit rendering methods like NeRF incur prohibitive computational costs. To overcome these limitations, the authors propose a lightweight 3D talking-face compression framework that uniquely integrates the FLAME parametric face model with 3D Gaussian Splatting (3DGS). By transmitting only compact facial metadata and leveraging compressed Gaussian attributes alongside optimized MLP weights for efficient reconstruction, the method drastically reduces bandwidth requirements. It achieves superior rate-distortion performance at extremely low bitrates compared to existing approaches, enabling high-quality, real-time 3D facial communication.