Score
Designs, implements, and analyzes end-to-end media processing pipelines for capture, preprocessing, encoding/decoding, transport and playback—covering audio, video and signal-processing paths and real-time streaming scenarios. Builds and tunes codecs and encoders, defines and applies audio quality metrics and QA/QC processes, and engineers low-latency, scalable data pipelines for audio capture/playback and signal processing.
This paper addresses the end-to-end Quality of Experience (QoE) assurance challenge for video streaming over best-effort networks. It systematically analyzes bottlenecks across the full pipeline—from video acquisition and compression (H.264/HEVC/AV1), upload, transcoding, CDN scheduling, adaptive bitrate (ABR) decision-making, to playback. We propose the first unified end-to-end pipeline analytical framework, classifying and modeling over 200 works along two orthogonal dimensions: methodology (heuristic, optimization, machine learning) and technical characteristics (codecs, super-resolution, etc.). The resulting methodology map is the most comprehensive to date, rigorously delineating performance boundaries and industrial deployment constraints for each approach. Our analysis identifies critical evolutionary trends—including ultra-low-latency live streaming, AI-native video coding, and edge-coordinated delivery—providing a systematic reference for both academic research and industry implementation.
To address challenges in distributed orchestration, lack of standardization, and insufficient elasticity for multimedia workflows in cloud–edge collaborative environments, this paper proposes and implements a cloud-native networked multimedia workflow system. The system introduces the first open-source implementation supporting the ISO/IEC 23090-8 NBMP international standard, and adopts a declarative microservice architecture built atop Kubernetes to uniformly abstract end-to-end media processing tasks—including ingestion, transcoding, packaging, and distribution. Leveraging NBMP-standardized APIs and containerized resource scheduling, it enables dynamic cross-environment deployment, fine-grained elastic scaling, and load balancing across cloud and edge infrastructures. Experimental evaluation in a real-world hybrid cloud–edge setting demonstrates sub-500 ms end-to-end latency for live streaming, a 40% throughput improvement, and strong scalability and engineering deployability.
Video streaming faces a fundamental trade-off between high visual quality and low latency. Conventional codecs, lacking contextual modeling capabilities, transmit all frame data, resulting in significant bandwidth redundancy. This paper proposes ELVIS, an end-to-end enhancement framework that introduces a novel collaborative paradigm: server-side rate-distortion-optimized encoding coupled with client-side generative reconstruction. Specifically, the server proactively discards redundant frames—e.g., low-motion regions—while the client reconstructs them using diffusion models or transformer-based generative techniques. ELVIS adopts a modular architecture, enabling flexible substitution of encoders, reconstruction models, and quality metrics (e.g., VMAF). Evaluated on standard benchmarks, ELVIS achieves a 11-point VMAF improvement and substantially reduces bandwidth requirements at equivalent perceptual quality. The framework establishes a scalable architectural foundation for integrating generative AI into real-time streaming systems.
To address the challenges of poor modularity, high latency, and limited scalability in real-time audio labeling tools for IP-based broadcast workflows, this paper proposes a lightweight audio tagging microservice architecture tailored for broadcast applications. The architecture leverages Docker containerization and RESTful API design, integrates pre-trained models (e.g., PANNs), and natively supports SMPTE ST 2110-30 for low-latency real-time audio stream analysis. It introduces a novel pluggable architecture that simultaneously ensures IP network compatibility and end-to-end real-time performance, significantly enhancing system flexibility and vendor interoperability. Experimental evaluation demonstrates an end-to-end latency under 200 ms and noise event detection accuracy exceeding 92% in live news and music broadcasting scenarios. These results validate the architecture’s adaptability and practicality across broadcast workflows—from small-scale production environments to large enterprise deployments.
High end-to-end latency (>5 seconds) in real-time media streaming—such as live sports, news, surveillance, and linear TV—remains a critical bottleneck. Method: This paper systematically surveys the evolution of low-latency live streaming in the IP era, proposing the “Live Latency Continuum” model to unify and characterize latency coupling across the entire pipeline—from acquisition to playback. It comprehensively analyzes HTTP-based adaptive streaming extensions (e.g., LL-DASH, LL-HLS), integrates QUIC/HTTP/3 transport optimizations, and incorporates client-side adaptive buffering strategies, enabling rigorous latency modeling, root-cause attribution, and engineering validation. Contribution/Results: The study demonstrates the technical feasibility of sub-one-second end-to-end latency and delivers a reusable, modular low-latency optimization framework. This work establishes a critical benchmark for both academic research and industrial standardization efforts in ultra-low-latency streaming.
This work addresses the lack of a unified, high-quality transcoding method for spatial audio across diverse acquisition formats—such as Ambisonics or microphone arrays—and arbitrary playback systems. The authors propose a general parametric framework that estimates spatial metadata of primary sources and ambient sound in the time–frequency domain, constructs a spatial covariance model tailored to the target playback setup, and derives an optimal linear downmix matrix. This approach supports independent rotation between acquisition and playback geometries and, for the first time, unifies processing for both Ambisonics and raw microphone array inputs. It accommodates arbitrary array configurations, variable numbers of sources, and arbitrary angular power distributions of ambient sound. Listening tests demonstrate that the method significantly outperforms existing parametric renderers across various content types and playback configurations, with particularly notable perceptual improvements for low-order or geometrically constrained arrays.
This work addresses the challenge of real-time streaming audiovisual character generation by simultaneously ensuring speech–text alignment, cross-segment visual consistency, and low-latency constraints. The authors propose a decoupled architecture comprising an LLM-driven coordinator that produces frame-level aligned audio conditions and employs a progress-aware pointer to maintain text–speech synchronization. A joint audiovisual DiT model performs localized bidirectional denoising within short temporal windows, augmented with a sink-token memory mechanism to suppress visual drift. Efficient deployment is achieved through a two-stage distillation strategy. Evaluated on a single H100 GPU, the method achieves real-time performance and outperforms existing baselines in text fidelity, audiovisual synchronization, visual quality, and streaming stability.
This work addresses playback stuttering in real-time streaming video generation caused by generation lag by proposing a dynamic serving mechanism that uses playout slack as a unified scheduling signal. The mechanism employs a three-level priority queue, re-homing, elastic sequence parallelism, and dual-mode Pareto routing to enable cross-stream resource reallocation and block-level quality adaptation, balancing timeliness and user experience without compromising generation quality. Experimental results on a 16-GPU H100 cluster demonstrate that the system improves continuous playback rate by 1.64–3.29× and reduces time-to-first-block by 1.61–9.65× compared to baseline approaches.
本文针对实时互动世界的流式音视频生成问题,提出了StreamAV-Bench基准,通过统一评估框架和32个细粒度维度的测试案例来评估13个代表性系统。
This study addresses the capability degradation caused by sequentially training streaming and few-step sampling for real-time audio-visual generation. To overcome this, we propose parallel training of causal and few-step adapters on a frozen diffusion Transformer backbone. Exploiting the approximate orthogonality of their update directions, the adapters are combined via direct addition based on model merging principles, thereby avoiding the interference inherent in chained fine-tuning without requiring explicit constraints. By integrating block-autoregressive attention with LoRA, the proposed method achieves continuous real-time generation at approximately 26 fps at 480×832 resolution. The resulting image quality is comparable to that of bidirectional teacher models and significantly surpasses conventional baselines.