Score
Designs and builds low-latency end-to-end media processing pipelines that capture, encode/decode, packetize, transport, synchronize, and render real‑time audio and video streams while handling buffering, jitter, and latency constraints. Integrates and configures frameworks and transports such as GStreamer, WebRTC, and WebSockets to implement signaling, transport (including NAT traversal and ICE), media tracks and data channels, and to connect pipelines to application logic.
High end-to-end latency (>5 seconds) in real-time media streaming—such as live sports, news, surveillance, and linear TV—remains a critical bottleneck. Method: This paper systematically surveys the evolution of low-latency live streaming in the IP era, proposing the “Live Latency Continuum” model to unify and characterize latency coupling across the entire pipeline—from acquisition to playback. It comprehensively analyzes HTTP-based adaptive streaming extensions (e.g., LL-DASH, LL-HLS), integrates QUIC/HTTP/3 transport optimizations, and incorporates client-side adaptive buffering strategies, enabling rigorous latency modeling, root-cause attribution, and engineering validation. Contribution/Results: The study demonstrates the technical feasibility of sub-one-second end-to-end latency and delivers a reusable, modular low-latency optimization framework. This work establishes a critical benchmark for both academic research and industrial standardization efforts in ultra-low-latency streaming.
This paper addresses the end-to-end Quality of Experience (QoE) assurance challenge for video streaming over best-effort networks. It systematically analyzes bottlenecks across the full pipeline—from video acquisition and compression (H.264/HEVC/AV1), upload, transcoding, CDN scheduling, adaptive bitrate (ABR) decision-making, to playback. We propose the first unified end-to-end pipeline analytical framework, classifying and modeling over 200 works along two orthogonal dimensions: methodology (heuristic, optimization, machine learning) and technical characteristics (codecs, super-resolution, etc.). The resulting methodology map is the most comprehensive to date, rigorously delineating performance boundaries and industrial deployment constraints for each approach. Our analysis identifies critical evolutionary trends—including ultra-low-latency live streaming, AI-native video coding, and edge-coordinated delivery—providing a systematic reference for both academic research and industry implementation.
This work addresses the persistent stuttering in low-latency live streaming, which occurs even when the encoded bitrate remains below available bandwidth, due to the inability of traditional packet-level congestion control to accurately estimate available bandwidth amid video frame encoding fluctuations. To resolve this, the authors propose Camel—the first frame-level congestion control algorithm tailored for low-latency live streaming—that decouples encoding-induced variability from bandwidth estimation using frame-level network feedback and introduces a burst-length control mechanism to dynamically optimize both average sending rate and burst patterns. The system comprises three core modules: a bandwidth/delay estimator, a congestion detector, and a burst controller. Deployed on a platform serving hundreds of millions of users, Camel increased 1080p stream share by 70.8%, raised media bitrate by 14.4%, and reduced stuttering by 14.1%; simulations further demonstrated up to 93.0% stutter reduction and a 23.9% improvement in bandwidth estimation accuracy.
This work proposes video-io, a generic Web component built upon WebRTC and Web Components standards, to address the limitations of existing WebRTC-based video communication services, which are typically confined to call- or room-centric models and tightly coupled with vendor-specific logic. By introducing named media stream abstractions, video-io keeps application logic on the client side and delegates vendor-specific signaling and access control to pluggable service connectors. This architecture decouples applications from underlying platforms, enabling write-once, deploy-anywhere capabilities across diverse environments. The authors demonstrate the framework’s cross-platform compatibility and architectural flexibility by implementing connectors for ten distinct systems, confirming its effectiveness in supporting flexible media publishing and subscription scenarios beyond traditional conferencing paradigms.
This work addresses the high latency and error accumulation inherent in traditional cascaded systems for real-time audiovisual interaction, which stem from modular fragmentation. To overcome these limitations, the authors propose the first end-to-end native streaming multimodal foundation model that unifies language, audio, and video into an interleaved token sequence processed within a single Transformer architecture, enabling full-duplex interaction. The system employs a causal encoder–decoder, block-wise causal attention, and a low-latency multimodal token scheduling strategy to support incremental processing at 160 ms granularity (25 fps) without relying on external speech recognition, synthesis, or animation modules. Empirical evaluation demonstrates an on-device response latency of approximately 200 ms; combined with 350 ms network latency, the total interaction latency reaches about 550 ms, marking the first achievement of sub-second full-duplex audiovisual communication.
This work proposes a fully decentralized web-based multiparty video conferencing system that eliminates the need for centralized media servers, thereby addressing the high costs and architectural complexity inherent in traditional approaches. By leveraging only lightweight notification services—such as email or push notifications—for signaling, the system establishes peer-to-peer (P2P) connections directly between clients to transmit audio, video, screen sharing, and text messages. Built upon WebRTC data channels and media streams, and implemented via browser extensions or Progressive Web Apps (PWAs), the solution significantly reduces deployment and operational overhead. This approach offers an efficient and practical decentralized conferencing alternative, particularly well-suited for resource-constrained environments.
Video streaming faces a fundamental trade-off between high visual quality and low latency. Conventional codecs, lacking contextual modeling capabilities, transmit all frame data, resulting in significant bandwidth redundancy. This paper proposes ELVIS, an end-to-end enhancement framework that introduces a novel collaborative paradigm: server-side rate-distortion-optimized encoding coupled with client-side generative reconstruction. Specifically, the server proactively discards redundant frames—e.g., low-motion regions—while the client reconstructs them using diffusion models or transformer-based generative techniques. ELVIS adopts a modular architecture, enabling flexible substitution of encoders, reconstruction models, and quality metrics (e.g., VMAF). Evaluated on standard benchmarks, ELVIS achieves a 11-point VMAF improvement and substantially reduces bandwidth requirements at equivalent perceptual quality. The framework establishes a scalable architectural foundation for integrating generative AI into real-time streaming systems.
This work addresses playback stuttering in real-time streaming video generation caused by generation lag by proposing a dynamic serving mechanism that uses playout slack as a unified scheduling signal. The mechanism employs a three-level priority queue, re-homing, elastic sequence parallelism, and dual-mode Pareto routing to enable cross-stream resource reallocation and block-level quality adaptation, balancing timeliness and user experience without compromising generation quality. Experimental results on a 16-GPU H100 cluster demonstrate that the system improves continuous playback rate by 1.64–3.29× and reduces time-to-first-block by 1.61–9.65× compared to baseline approaches.
This work proposes a thinker-performer architecture to enhance the output resolution of audio-visual interactive models under low-latency constraints. The thinker module, running efficiently on a single GPU, handles language understanding and state maintenance, while the performer module—deployed across multiple GPUs using Ulysses parallelism—specializes in generating high-resolution visual streams, scaling from 192×336 to 640×368. By injecting language states as key/value (K/V) conditions into the performer and employing techniques such as K/V cache presharding, latent sequence sharding, and denoising with aggregation, the design eliminates cross-GPU transmission of language sequences, substantially improving computational efficiency. The system achieves approximately 200 ms model latency at 25 FPS and an end-to-end remote interaction latency of around 550 ms, enabling real-time embodied interaction with clear, discernible mid-shot representations of pose, gestures, and scene layout.
本文提出了一种基于边缘-云的实时视频理解系统,集成了多种视觉-语言模型后端,并通过WebRTC实现了低延迟交互。