Score
Designs, implements, and analyzes audio and video codec algorithms, compression techniques, and encoding/decoding pipelines to improve rate‑distortion efficiency, computational cost, latency, and perceptual quality; this includes work on standard-compliant video/audio codecs, codec integration into playback/streaming systems, and both algorithmic and implementation-level optimizations.
Conventional evaluation of learned video compression via direct averaging of rate-distortion (RD) curves across test sequences introduces systematic bias—outlier sequences disproportionately influence the mean curve, obscuring true codec performance on the majority of sequences. Method: The authors systematically analyze this issue through analytical modeling, empirical validation on the UVG dataset, and comparative BD-rate analysis. Contribution/Results: They demonstrate that rankings derived from averaged RD curves frequently contradict those obtained from per-sequence metric averaging (e.g., PSNR or MS-SSIM), revealing fundamental inconsistencies in current practice. The paper advocates a return to the traditional video coding standard: per-sequence RD analysis followed by arithmetic averaging of distortion metrics at fixed bitrates (or BD-rate). Empirical results confirm this approach yields more robust, consistent, and statistically reliable performance assessments. This work establishes both theoretical justification and practical guidelines for redefining evaluation paradigms in learned video compression.
This work addresses the challenge of balancing compression efficiency and computational complexity in practical deployments of intelligent video coding. We propose an end-to-end low-complexity coding framework tailored to standardized common test conditions, integrated into the AVS-EEM platform. By leveraging a customized neural network architecture, efficient training strategies, and inference optimization techniques—all while strictly adhering to conventional coding common test conditions—the proposed approach substantially reduces computational overhead. After more than two years of iterative development, the latest model significantly outperforms the AVS3 reference software in compression performance under identical test conditions, marking a critical step toward the standardization and practical adoption of end-to-end intelligent video coding.
This study addresses the challenge of balancing compression efficiency and perceptual audio quality in audio codec selection. Methodologically, it introduces a human auditory perception–centered evaluation framework integrating the Perceptual Evaluation of Audio Quality (PEAQ) objective model, multi-bitrate encoding performance testing, time-frequency spectrogram visualization, and multidimensional quality analysis to quantitatively characterize distortion mechanisms affecting perceived sound quality. Its key contribution lies in the first systematic, cross-codec comparison—under standardized experimental conditions—of mainstream codecs (e.g., MP3, AAC, Opus, FLAC) along their rate–perceptual-quality trade-off curves, revealing distinct patterns of perceptual degradation. The results provide reproducible empirical evidence and application-oriented, quality-efficiency co-optimization guidelines for codec selection across diverse use cases.
To address the low sub-pixel motion compensation accuracy, high computational overhead, and inferior compression performance of learned video codecs relative to HEVC/VVC, this paper proposes three synergistic optimizations: (1) replacing bilinear interpolation with a learnable high-order interpolation filter; (2) parameterizing motion information at the block level to reduce motion field redundancy; and (3) introducing a finite-precision motion vector modeling mechanism to minimize quantization error while preserving compensation accuracy. Evaluated within the COOL-CHIC framework, the proposed method achieves an average BD-rate reduction of 10.2% and reduces motion-compensation-related decoding computation from 391 to 214 MACs per pixel—a 45.3% decrease—significantly narrowing the performance gap with conventional codecs. The implementation is publicly available.
This paper addresses the end-to-end Quality of Experience (QoE) assurance challenge for video streaming over best-effort networks. It systematically analyzes bottlenecks across the full pipeline—from video acquisition and compression (H.264/HEVC/AV1), upload, transcoding, CDN scheduling, adaptive bitrate (ABR) decision-making, to playback. We propose the first unified end-to-end pipeline analytical framework, classifying and modeling over 200 works along two orthogonal dimensions: methodology (heuristic, optimization, machine learning) and technical characteristics (codecs, super-resolution, etc.). The resulting methodology map is the most comprehensive to date, rigorously delineating performance boundaries and industrial deployment constraints for each approach. Our analysis identifies critical evolutionary trends—including ultra-low-latency live streaming, AI-native video coding, and edge-coordinated delivery—providing a systematic reference for both academic research and industry implementation.
This work addresses the challenges of deploying deep learning-based in-loop filtering in consumer electronics—namely high computational complexity, limited memory bandwidth, and stringent power constraints—by proposing the first hardware-oriented 3D classification framework. It systematically surveys deep learning filtering approaches in video coding, covering integration strategies, exploitation of coding-side information, and lightweight network design. In alignment with recent standardization efforts in JVET’s Neural Network Video Coding (NNVC), the study provides an in-depth analysis of the trade-offs between rate-distortion performance and hardware feasibility. It delineates a clear evolutionary pathway from high-performance models toward low-power, real-time NPU-friendly architectures, identifies critical challenges such as inference latency and error propagation, and offers both theoretical insights and a practical roadmap for deploying intelligent video coding on edge devices.
Existing learned image codecs struggle to simultaneously achieve high perceptual quality and real-time performance on edge devices. This work systematically investigates key modeling choices affecting practicality and proposes a unified optimization framework that integrates differentiable compression, perception-driven loss optimization, and ablation-guided module design. For the first time, it conducts a large-scale neural architecture search (NAS) over millions of backbone configurations under explicit latency and quality constraints. The resulting efficient codec achieves 230 ms encoding and 150 ms decoding for 12MP images on an iPhone 17 Pro Max. In subjective evaluations, it reduces bitrate by 2.3–3× compared to AV1/VVC and further improves upon state-of-the-art learned methods by 20–40% in bitrate savings.
This work addresses the challenges of high computational overhead and real-time processing in video stream analysis for vision-language model services. The authors propose an end-to-end online optimization framework that, for the first time, leverages metadata naturally generated during video decoding as a low-cost runtime signal to jointly guide video decoding, Vision Transformer (ViT) patch pruning, and selective refresh of large language model key-value (KV) caches—without requiring offline training. By integrating metadata-driven online patch pruning, selective KV cache updates, and compressed bitstream passthrough, the method achieves up to 3× higher throughput and an 87% reduction in GPU compute cost compared to the best existing baseline, while maintaining F1 scores within only 0–8% degradation.
This work addresses the challenge of maintaining perceptual quality consistency in learned video codecs when deployed across varying spatial resolutions, a scenario that typically necessitates retraining or rate-distortion parameter tuning. Building upon the MS-VQ-VAE framework, the study systematically investigates the impact of codebook capacity and spatial resolution on perceptual quality using the UCF101 dataset. The findings reveal that codebook capacity exerts an influence approximately ten times greater than that of resolution, with higher resolutions yielding superior entropy efficiency. These insights offer a novel perspective for designing discrete tokenizers in multi-resolution video compression and generative models. Experimental results demonstrate that the proposed method achieves LPIPS scores surpassing H.264 by 25–52% at 128×128 resolution and outperforming H.265 by 21–37% at 256×256, all while operating at comparable or lower bitrates.
General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.