Score
Designs and implements end-to-end systems that convert, encode, decode, and package audio and video between formats by composing and orchestrating transcoding pipelines and format-transcoding steps. Builds, integrates with, and analyzes ffmpeg-based tooling and internals to automate, optimize, debug, and manage command-line or programmatic transcoding workflows, resource usage, and format compatibility.
This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.
Instruction-tuned large language models for code often lack boundary awareness in Fill-in-the-Middle (FIM) tasks, necessitating post-hoc truncation to discard extraneous output—yet cross-lingual truncation strategies are inconsistent and suboptimal. Method: The authors conduct the first systematic investigation into whether supervised fine-tuning (SFT) can inherently improve contextual and boundary alignment, eliminating reliance on heuristic post-processing. Using the Qwen2.5-Coder family, they introduce a binary evaluation paradigm—“complete-line vs. random-fragment”—on HumanEval Infilling and SAFIM benchmarks. Contribution/Results: SFT significantly enhances boundary-aware generation: fine-tuned models achieve optimal performance without post-processing in complete-line scenarios, while truncation remains necessary only for random fragments. This reveals the conditional necessity of post-processing in FIM, challenging the assumption of universal truncation requirements and providing empirical grounding for boundary-aware code generation.
This work addresses the challenge of balancing compression efficiency and computational complexity in practical deployments of intelligent video coding. We propose an end-to-end low-complexity coding framework tailored to standardized common test conditions, integrated into the AVS-EEM platform. By leveraging a customized neural network architecture, efficient training strategies, and inference optimization techniques—all while strictly adhering to conventional coding common test conditions—the proposed approach substantially reduces computational overhead. After more than two years of iterative development, the latest model significantly outperforms the AVS3 reference software in compression performance under identical test conditions, marking a critical step toward the standardization and practical adoption of end-to-end intelligent video coding.
This paper addresses the end-to-end Quality of Experience (QoE) assurance challenge for video streaming over best-effort networks. It systematically analyzes bottlenecks across the full pipeline—from video acquisition and compression (H.264/HEVC/AV1), upload, transcoding, CDN scheduling, adaptive bitrate (ABR) decision-making, to playback. We propose the first unified end-to-end pipeline analytical framework, classifying and modeling over 200 works along two orthogonal dimensions: methodology (heuristic, optimization, machine learning) and technical characteristics (codecs, super-resolution, etc.). The resulting methodology map is the most comprehensive to date, rigorously delineating performance boundaries and industrial deployment constraints for each approach. Our analysis identifies critical evolutionary trends—including ultra-low-latency live streaming, AI-native video coding, and edge-coordinated delivery—providing a systematic reference for both academic research and industry implementation.
This work addresses cross-modal video-to-audio generation with emphasis on semantic consistency and frame-level temporal alignment. To this end, we propose an alignment-aware framework featuring: (i) a lightweight visual encoder for efficient video representation extraction; (ii) learnable auxiliary embeddings that explicitly model audio–video correspondence; and (iii) multi-scale temporal data augmentation coupled with end-to-end joint training to enforce temporal coherence. Our key innovation lies in an implicit alignment mechanism, which reveals the critical role of auxiliary embeddings and augmentation strategies in achieving precise synchronization. We further introduce the first comprehensive evaluation paradigm specifically designed for audio–video alignment. Experiments demonstrate state-of-the-art performance in both audio fidelity—measured by STFT-L1 and PESQ—and frame-level synchronization accuracy—quantified by SyncScore—establishing a new benchmark for photorealistic audiovisual generation.
This work addresses the inefficiencies of the edit-compile-reload cycle in audio DSP development and the loss of parameter bindings caused by structural changes in DSP code. To resolve these issues, the authors propose a dual-mode CLAP plugin compilation system for the Faust language that integrates static ahead-of-time (AOT) compilation with dynamic interpreted execution. The system introduces an innovative address-based parameter identity matching algorithm and a stable slot allocation mechanism, enabling runtime hot reloading while preserving consistent parameter identities and host automation bindings. As the first officially integrated Faust-to-CLAP compilation pathway, the implementation achieves significant gains in development iteration speed and parameter stability with only approximately 2,400 lines of C++/Python code.
Existing feature attribution methods incur substantial computational and evaluation overhead due to the inclusion of numerous irrelevant cross-layer transformer (CLT) features introduced by uniform sampling. This work proposes PIE, a novel framework that establishes the first native end-to-end interpretability pipeline for CLTs, seamlessly integrating pruning, automated explanation, and evaluation. Central to PIE are two new mechanisms: Feature Attribution Patching (FAP) and FAP-Synergy, a synergy-aware re-ranking strategy. Evaluated on the IOI and Doc-String tasks under strict computational budgets, FAP-based methods significantly enhance behavioral fidelity and explanation quality—achieving performance comparable to that of randomly selected sets of approximately 4,000 features using only K=100, thereby yielding a ~40× compression ratio and substantially reducing interpretability costs.
This work addresses the lack of a unified, high-quality transcoding method for spatial audio across diverse acquisition formats—such as Ambisonics or microphone arrays—and arbitrary playback systems. The authors propose a general parametric framework that estimates spatial metadata of primary sources and ambient sound in the time–frequency domain, constructs a spatial covariance model tailored to the target playback setup, and derives an optimal linear downmix matrix. This approach supports independent rotation between acquisition and playback geometries and, for the first time, unifies processing for both Ambisonics and raw microphone array inputs. It accommodates arbitrary array configurations, variable numbers of sources, and arbitrary angular power distributions of ambient sound. Listening tests demonstrate that the method significantly outperforms existing parametric renderers across various content types and playback configurations, with particularly notable perceptual improvements for low-order or geometrically constrained arrays.
This work addresses the challenge of real-time streaming audiovisual character generation by simultaneously ensuring speech–text alignment, cross-segment visual consistency, and low-latency constraints. The authors propose a decoupled architecture comprising an LLM-driven coordinator that produces frame-level aligned audio conditions and employs a progress-aware pointer to maintain text–speech synchronization. A joint audiovisual DiT model performs localized bidirectional denoising within short temporal windows, augmented with a sink-token memory mechanism to suppress visual drift. Efficient deployment is achieved through a two-stage distillation strategy. Evaluated on a single H100 GPU, the method achieves real-time performance and outperforms existing baselines in text fidelity, audiovisual synchronization, visual quality, and streaming stability.