Score
Designs and implements the integration of compression codecs into software or hardware pipelines, including adding or adapting internal codec modules (e.g., motion-estimation) and correctly connecting inter-coding stages. Builds evaluation and benchmarking protocols to measure and analyze bitrate, latency, and quality trade-offs and to compare integrated codecs against existing implementations.
This work addresses the challenge of balancing compression efficiency and computational complexity in practical deployments of intelligent video coding. We propose an end-to-end low-complexity coding framework tailored to standardized common test conditions, integrated into the AVS-EEM platform. By leveraging a customized neural network architecture, efficient training strategies, and inference optimization techniques—all while strictly adhering to conventional coding common test conditions—the proposed approach substantially reduces computational overhead. After more than two years of iterative development, the latest model significantly outperforms the AVS3 reference software in compression performance under identical test conditions, marking a critical step toward the standardization and practical adoption of end-to-end intelligent video coding.
To address scalability and ultra-low-latency transmission challenges in real-time audio-video conferencing systems on FPGAs, this paper proposes a fully hardware-coordinated architecture. The design implements M-JPEG video encoding, PCM audio sampling, and a lightweight UDP protocol stack entirely in SystemVerilog, and integrates them end-to-end on a Nexys4 DDR FPGA. A key innovation is the tight hardware-level coupling between audio-video encoding and network transmission, enabling stable 30 FPS video streaming with synchronized decoding. Modular verification employs Cocotb, synthesis and implementation are performed in Vivado, and host-side synchronization playback is achieved via Python-based parsing. Experimental results demonstrate an end-to-end latency under 120 ms and throughput sufficient for real-time communication. The system is rigorously validated through both simulation and hardware deployment, confirming high stability, scalability, and suitability for resource-constrained FPGA platforms.
Video compression reduces transmission overhead but degrades visual quality, leading to significant performance drops in downstream vision tasks (e.g., object detection, action recognition). To address this, we propose a codec-aware, plug-and-play video enhancement framework that requires no modification to existing codecs and introduces zero inference latency. Our approach features two key innovations: (1) a novel hierarchical compressed-domain awareness mechanism that jointly models spatiotemporal priors from bitstreams; and (2) co-optimized lightweight frame-level enhancement (BAE) and compression-aware adaptation (CAA) networks, enabling zero-intrusion, cross-standard adaptability across H.264, H.265, and AV1. Evaluated on multiple benchmarks, our method consistently outperforms state-of-the-art approaches, improving average downstream task accuracy by 3.2–5.7% while maintaining real-time deployment capability.
To address bandwidth constraints and underutilized compression efficiency in edge video analytics, this paper proposes the first macroblock-level adaptive learning compression framework tailored for modern block-based encoders (e.g., H.264). Our method introduces the first deep learning–based prediction of macroblock-level quantization parameters, enabling fine-grained, end-to-end joint optimization of bitrate and analytical accuracy, while seamlessly integrating into existing edge analytics pipelines. The core innovation lies in explicitly modeling task-specific analytical requirements as quality control objectives, thereby achieving optimal bit allocation under analytical accuracy constraints. Experiments demonstrate that, while preserving target detection and recognition accuracy, our approach reduces bitrate by 38.7% on average—up to 50.4%—yielding a 3.01× improvement in compression efficiency over conventional methods.
To address the low sub-pixel motion compensation accuracy, high computational overhead, and inferior compression performance of learned video codecs relative to HEVC/VVC, this paper proposes three synergistic optimizations: (1) replacing bilinear interpolation with a learnable high-order interpolation filter; (2) parameterizing motion information at the block level to reduce motion field redundancy; and (3) introducing a finite-precision motion vector modeling mechanism to minimize quantization error while preserving compensation accuracy. Evaluated within the COOL-CHIC framework, the proposed method achieves an average BD-rate reduction of 10.2% and reduces motion-compensation-related decoding computation from 391 to 214 MACs per pixel—a 45.3% decrease—significantly narrowing the performance gap with conventional codecs. The implementation is publicly available.
为提高压缩视频质量,提出MDFI方法,通过多域特征融合策略和新颖的帧预测特征转换模块处理预测信息,有效结合时空特性、跨频表示和压缩域预测信息。
研究使用大型语言模型设计视频编码工具,通过生成-评估循环改进Planar模式,实验表明该方法可提高编码效率。
This work addresses the limitation of existing image compression methods that neglect the joint optimization of statistical and semantic information in entropy models when adapting pretrained codecs, thereby constraining the effectiveness of parameter-efficient fine-tuning. To overcome this, the authors propose S2-CoT, a structure–semantics co-tuning framework that systematically analyzes and coordinates adapter type and placement. Specifically, they introduce a Structure-Fidelity Adapter (SFA) for the codec and a Semantic Context Adapter (SCA) for the entropy model, enabling dual-adapter joint optimization through parameter-efficient fine-tuning, spatial–frequency feature fusion, and channel-wise context modeling. Evaluated on four mainstream codecs, S2-CoT achieves performance comparable to full fine-tuning using only a minimal number of trainable parameters, significantly enhancing compression efficiency for machine vision tasks and establishing new state-of-the-art results.
本文提出了一种名为ApproxSSIMate的方法,通过从PSNR估计SSIM并结合序列统计信息,来解决在编码过程中高效评估感知质量的问题。
This work addresses the storage, transmission, and deployment challenges posed by the massive parameter counts of large language models by introducing, for the first time, a systematic application of modern video compression techniques to model weight quantization. The proposed method integrates affine quantization with advanced video coding standards such as VVC/H.266, naturally aligning with the structural properties of weight matrices without requiring fine-tuning or calibration data. It demonstrates strong generalization across diverse tensor types. Experimental results on the LLaMA-3-8B model at 2-bit compression show a more than 1.5× reduction in perplexity and a 21% improvement in downstream task accuracy compared to existing approaches, substantiating the method’s efficiency, robustness, and broad applicability.