Score
Designs, implements, and evaluates systems and components that enable live (real‑time) streaming of audio, video, and associated interactive data over networks. This includes capture and encoding pipelines, streaming protocols and delivery/CDN architecture for low‑latency transmission, client playback and interaction UX (chat, reactions, synchronization), content moderation and rights handling, and real‑time monitoring and analytics.
High end-to-end latency (>5 seconds) in real-time media streaming—such as live sports, news, surveillance, and linear TV—remains a critical bottleneck. Method: This paper systematically surveys the evolution of low-latency live streaming in the IP era, proposing the “Live Latency Continuum” model to unify and characterize latency coupling across the entire pipeline—from acquisition to playback. It comprehensively analyzes HTTP-based adaptive streaming extensions (e.g., LL-DASH, LL-HLS), integrates QUIC/HTTP/3 transport optimizations, and incorporates client-side adaptive buffering strategies, enabling rigorous latency modeling, root-cause attribution, and engineering validation. Contribution/Results: The study demonstrates the technical feasibility of sub-one-second end-to-end latency and delivers a reusable, modular low-latency optimization framework. This work establishes a critical benchmark for both academic research and industrial standardization efforts in ultra-low-latency streaming.
To address the challenge of balancing low latency and high safety in real-time live video content moderation, this paper proposes an end-to-end dynamic filtering framework tailored for social platforms. The method decentralizes content filtering to clients for the first time, leveraging an extended Media over QUIC (MoQ) protocol to enable GOP-level streaming policy enforcement, collaborative client-side analysis, and lightweight photosensitive seizure detection. The system selectively removes only non-compliant segments while immediately resuming playback—ensuring safety for photosensitive users with an end-to-end latency of only 200–500 ms (i.e., one GOP duration). Experimental results demonstrate substantial improvements over conventional centralized moderation: the framework achieves breakthroughs in accessibility, real-time performance, and safety, establishing a new trade-off frontier between latency and security in live streaming moderation.
This paper addresses the end-to-end Quality of Experience (QoE) assurance challenge for video streaming over best-effort networks. It systematically analyzes bottlenecks across the full pipeline—from video acquisition and compression (H.264/HEVC/AV1), upload, transcoding, CDN scheduling, adaptive bitrate (ABR) decision-making, to playback. We propose the first unified end-to-end pipeline analytical framework, classifying and modeling over 200 works along two orthogonal dimensions: methodology (heuristic, optimization, machine learning) and technical characteristics (codecs, super-resolution, etc.). The resulting methodology map is the most comprehensive to date, rigorously delineating performance boundaries and industrial deployment constraints for each approach. Our analysis identifies critical evolutionary trends—including ultra-low-latency live streaming, AI-native video coding, and edge-coordinated delivery—providing a systematic reference for both academic research and industry implementation.
This study systematically investigates the Quality-of-Experience (QoE) impact mechanisms of the QUIC protocol in multi-client video streaming scenarios. Focusing on two representative use cases—video-on-demand (VoD) and low-latency live (LLL) streaming—we employ a trace-driven simulation framework to analyze cross-layer interactions between mainstream QUIC implementations (featuring congestion control algorithms including Cubic and BBR) and adaptive bitrate (ABR) strategies (BOLA, Pensieve). We empirically reveal, for the first time, that identical congestion control algorithms exhibit substantial performance variation across different QUIC implementations, leading to measurable discrepancies in key QoE metrics—namely, stall ratio, startup latency, and average bitrate. Building upon this insight, we propose a QUIC–ABR cross-layer co-optimization framework that jointly adapts congestion window feedback and bitrate selection decisions. Our approach achieves quantifiable QoE improvements: a 32% reduction in stall ratio and an 18% increase in average video quality.
This work addresses the challenge of sustaining ultra-low-latency (millisecond-scale), reliable data delivery over dynamically fluctuating wireless networks—critical for remote collaboration and immersive AR/VR applications. We propose the first end-to-end data delivery architecture driven by real-world network measurements, integrating a software-defined protocol stack, synthetic channel modeling, adaptive forward error correction, latency-aware scheduling, and closed-loop channel state feedback. Experimental evaluation on live, time-varying wireless links demonstrates an end-to-end latency standard deviation ≤1.2 ms and a 99th-percentile latency <8 ms—substantially outperforming TCP and QUIC. Our key contribution is the first practical realization of millisecond-level latency stability, providing verifiable, low-overhead transport-layer support for high-fidelity immersive interaction.
This work proposes video-io, a generic Web component built upon WebRTC and Web Components standards, to address the limitations of existing WebRTC-based video communication services, which are typically confined to call- or room-centric models and tightly coupled with vendor-specific logic. By introducing named media stream abstractions, video-io keeps application logic on the client side and delegates vendor-specific signaling and access control to pluggable service connectors. This architecture decouples applications from underlying platforms, enabling write-once, deploy-anywhere capabilities across diverse environments. The authors demonstrate the framework’s cross-platform compatibility and architectural flexibility by implementing connectors for ten distinct systems, confirming its effectiveness in supporting flexible media publishing and subscription scenarios beyond traditional conferencing paradigms.
This work addresses the high latency and error accumulation inherent in traditional cascaded systems for real-time audiovisual interaction, which stem from modular fragmentation. To overcome these limitations, the authors propose the first end-to-end native streaming multimodal foundation model that unifies language, audio, and video into an interleaved token sequence processed within a single Transformer architecture, enabling full-duplex interaction. The system employs a causal encoder–decoder, block-wise causal attention, and a low-latency multimodal token scheduling strategy to support incremental processing at 160 ms granularity (25 fps) without relying on external speech recognition, synthesis, or animation modules. Empirical evaluation demonstrates an on-device response latency of approximately 200 ms; combined with 350 ms network latency, the total interaction latency reaches about 550 ms, marking the first achievement of sub-second full-duplex audiovisual communication.
This work addresses the persistent stuttering in low-latency live streaming, which occurs even when the encoded bitrate remains below available bandwidth, due to the inability of traditional packet-level congestion control to accurately estimate available bandwidth amid video frame encoding fluctuations. To resolve this, the authors propose Camel—the first frame-level congestion control algorithm tailored for low-latency live streaming—that decouples encoding-induced variability from bandwidth estimation using frame-level network feedback and introduces a burst-length control mechanism to dynamically optimize both average sending rate and burst patterns. The system comprises three core modules: a bandwidth/delay estimator, a congestion detector, and a burst controller. Deployed on a platform serving hundreds of millions of users, Camel increased 1080p stream share by 70.8%, raised media bitrate by 14.4%, and reduced stuttering by 14.1%; simulations further demonstrated up to 93.0% stutter reduction and a 23.9% improvement in bandwidth estimation accuracy.
This study addresses the critical issue of high end-to-end latency in real-time 3D voxel streaming, which severely degrades immersion, induces motion sickness, and impedes user interaction. The work presents the first hierarchical quantitative analysis of latency across the system stack, decomposing it into application, transport protocol, and network layers, and identifies bottleneck sources through empirical measurement. Building on these insights, the authors propose targeted system-level optimizations spanning protocol design, network scheduling, and rendering scheduling. Experimental evaluation demonstrates that the proposed approach substantially reduces end-to-end latency, thereby enhancing system responsiveness, scalability, and overall user experience.
This work addresses playback stuttering in real-time streaming video generation caused by generation lag by proposing a dynamic serving mechanism that uses playout slack as a unified scheduling signal. The mechanism employs a three-level priority queue, re-homing, elastic sequence parallelism, and dual-mode Pareto routing to enable cross-stream resource reallocation and block-level quality adaptation, balancing timeliness and user experience without compromising generation quality. Experimental results on a 16-GPU H100 cluster demonstrate that the system improves continuous playback rate by 1.64–3.29× and reduces time-to-first-block by 1.61–9.65× compared to baseline approaches.