Score
Designs, implements, and evaluates algorithms that partition sequential data—such as audio or text—into variable-sized segments (chunks) along axes like time, denoising steps, or other temporal dimensions, and that adapt chunk boundaries and sizes to constraints such as device computation or latency budgets. Builds chunking policies and mechanisms that support within-chunk autoregressive processing, cross-chunk bidirectional planning, and joint chunking across multiple axes, and analyzes tradeoffs among context window, computational cost, and generation or denoising quality.
This work addresses fundamental challenges in multimodal AI systems—namely, the absence of a unified chunking framework, cross-modal semantic inconsistency, and asynchronous information density coupled with noise interference. We propose the first comprehensive, modality-agnostic chunking taxonomy covering text, images, audio, video, and cross-modal data. Methodologically, we integrate fixed-size windowing, recursive text splitting, object-level visual chunking, silence-aware audio segmentation, and scene-aware video segmentation, implemented via a reusable technical framework leveraging LangChain, Detectron2, and PySceneDetect. Our contributions include: (1) a systematic characterization of the granularity–context trade-off; (2) a novel cross-modal chunking mechanism preserving semantic consistency; (3) identification and formal modeling of asynchronous information density and alignment noise as open problems; and (4) foundational theoretical and practical groundwork for adaptive, learning-driven, and task-specific chunking methodologies.
This work addresses the limited parallelizability of the classical dynamic time warping (DTW) algorithm, which suffers from quadratic time and memory complexity. The authors propose Segmental DTW, a novel approach that decomposes global sequence alignment into local subsequence DTW computations that can be executed in parallel, followed by a segment-level dynamic programming step to integrate the partial alignments. This method achieves near-full parallelism while preserving alignment accuracy comparable to standard DTW. Theoretical analysis and empirical evaluation on Chopin Mazurka audio alignment tasks demonstrate that one variant of the proposed method outperforms existing approaches in both computational efficiency and alignment performance.
This study addresses the inability of Large Audio-Language Models (LALMs) to perform media chapterization due to insufficient editorial judgment. We propose AudioChaps, a framework that innovatively bypasses supervised fine-tuning cold-start by directly aligning models via Group Relative Policy Optimization (GRPO) integrated with Chain-of-Thought (CoT) reasoning, supported by a specialized dataset designed to enhance editorial decision-making. Experimental results demonstrate that AudioChaps-R1 achieves a 49-point F1-score improvement over state-of-the-art baselines, successfully enabling the precise conversion of unstructured audio into navigable structured media. These findings significantly advance the practical utility of LALMs in real-world production workflows by effectively bridging the gap between raw audio processing and structured content organization through reinforcement learning-based alignment.
This work addresses the limitations of fixed-position chunking in discrete diffusion language models, which often disrupts semantic coherence and reduces modeling efficiency. The authors propose a dynamic semantic chunking mechanism that leverages differentiable Chunking Attention driven by learnable subspaces to cluster tokens into content-adaptive semantic blocks. Autoregressive diffusion denoising is then performed under block-wise causal masking. By replacing rigid positional chunks with semantically meaningful blocks, this approach enables end-to-end dynamic segmentation and strictly generalizes conventional block-based diffusion models. Experiments across model scales up to 1.5B parameters demonstrate that the method significantly outperforms both unstructured and position-based chunking baselines on downstream tasks, with performance gains emerging early in training and remaining consistent across model sizes.
Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Existing compression methods typically rely on fixed window partitioning and attention-based pruning, which overlook the piecewise semantic structure of audio-visual signals and become fragile under aggressive token reduction. We propose Dynamic Audio-driven Semantic cHunking (DASH), a training-free framework that aligns token compression with semantic structure. DASH treats audio embeddings as a semantic anchor and detects boundary candidates via cosine-similarity discontinuities, inducing dynamic, variable-length segments that approximate the underlying piecewise-coherent organization of the sequence. These boundaries are projected onto video tokens to establish explicit cross-modal segmentation. Within each segment, token retention is determined by a tri-signal importance estimator that fuses structural boundary cues, representational distinctiveness, and attention-based salience, mitigating the sparsity bias of attention-only selection. This structure-aware allocation preserves transition-critical tokens while reducing redundant regions. Extensive experiments on AVUT, VideoMME, and WorldSense demonstrate that DASH maintains superior accuracy while achieving higher compression ratios compared to prior methods. Code is available at: https://github.com/laychou666/DASH.
This work addresses the challenge of optimizing compression ratios in token-free hierarchical models with byte-level dynamic chunking by proposing an Adaptive Target Dynamic Chunking (ATDC) mechanism. ATDC introduces curriculum learning into dynamic chunking control for the first time, progressively increasing the target compression ratio from low to high during training to ensure stable optimization. The method models the evolution of chunking through Bytes Per Innermost Chunk (BPIC). Evaluated on the FineWeb-Edu 100B dataset, ATDC achieves bits-per-byte (BPB) performance comparable to both token-level and byte-level baselines while demonstrating markedly improved training stability and superior performance across multiple downstream tasks.
This work addresses the challenge that existing text-to-audio systems struggle to generate spoken audio with clear speech naturally integrated into ambient soundscapes, often suffering from muffled voices or insufficient temporal control due to reliance on post-processing. To overcome these limitations, the authors propose VoxAudio, a causal autoregressive flow-matching model featuring a chunk-wise causal factorization architecture that enables sliding-window streaming inference and leverages random-chunk pretraining for flexible generation at arbitrary granularities. Key innovations include multi-reward negative perceptual fine-tuning (NFT) for multi-objective preference optimization and the construction of VoxCorpus—the first speech-centric audio corpus with precise voice timing annotations—alongside the VoxBench evaluation benchmark. Experiments demonstrate that VoxAudio significantly outperforms current methods in semantic fidelity, linguistic accuracy, auditory aesthetics, and temporal alignment, while supporting efficient, variable-length streaming audio synthesis.
This study addresses the challenge of efficiently allocating computational resources under a fixed budget to optimize speech model performance. The authors develop a unified framework to systematically investigate the joint impact of model scale, input duration, and representation resolution on automatic speech recognition (ASR) and speech emotion recognition (SER). Through large-scale scaling experiments, efficient LoRA-based fine-tuning, and multi-granularity computational cost modeling, they uncover nonlinear scaling laws across these dimensions and propose design principles for identifying optimal operating points that balance performance and efficiency. Key findings include diminishing returns from increasing model size, peak SER performance at approximately 4 seconds of audio input, and the ability to significantly reduce encoder resolution with less than 3% performance degradation while yielding substantial computational savings.
This study addresses the challenges of uneven inference budget allocation and inconsistent guidance criteria in chunked sequence generation by proposing a budget-matched chunked guidance framework. Based on a two-dimensional lattice Feynman-Kac model, we design a sampler that optimizes resampling strategies across denoising steps and chunk indices. We derive a scaling theorem to establish design precision, enabling prefix-evaluable reward twisting without additional estimators, and formalize the theory of optimal resampling. Experimental results demonstrate that music-dance alignment improves from 0.234 to 0.441, with prompt adherence reaching 0.560. These outcomes significantly surpass Best-of-N baselines and are further corroborated by human preference evaluations.
This work addresses the structural conflicts arising from heterogeneous audio components—speech, music, and sound effects—in unified audio generation, which challenge the use of a shared backbone and demand fine-grained adaptability within individual segments. To this end, the authors propose SonicWeave, a flow-matching-based unified generative model featuring a Chunk-wise Prior-Evidence Mixture-of-Experts (CPE-MoE) mechanism. Its core innovation is a conflict-gated prior-evidence routing strategy that dynamically balances global textual conditioning and local contextual cues based on the reliability of local acoustic states during the diffusion process, enabling temporally coherent expert dispatching. Experiments demonstrate that SonicWeave outperforms both dense and vanilla MoE baselines across TTS, TTA, and TTM benchmarks, with superior performance in complex compositional scenarios; routing analysis further reveals content-dependent specialization of experts across diffusion stages.