Score
Designs and implements decoders that generate multiple candidate trajectories by iteratively producing and refining trajectory segments or full rollouts—often running segmental or parallel rollouts and recurrent decoding steps—to improve long‑horizon prediction accuracy. Builds algorithms for managing, selecting, combining, or reranking trajectory subsets across decoding iterations to capture segment‑wise temporal dependencies and stabilize multi‑trajectory outputs.
Autoregressive models suffer from high real-time inference latency due to sequential dependency, and conventional compression techniques—such as pruning and quantization—often incur significant accuracy degradation. To address this, we propose a unified generation–refinement decoding framework. Methodologically, we first establish a taxonomy of generation strategies—including n-gram matching and draft-model-based approaches—and refinement mechanisms—spanning single-step verification and iterative optimization. The framework integrates speculative decoding, multi-round verification, knowledge distillation from draft models, and hardware-aware scheduling for efficient deployment across heterogeneous platforms. Evaluated on text, image, and speech generation tasks, our approach achieves an average 2.1× speedup in end-to-end latency while sustaining minimal accuracy loss (<0.5% in BLEU, CLIP, and FID metrics). This work provides both scalable theoretical foundations and system-level implementation strategies for real-time large language and multimodal model applications.
This work addresses the computational bottleneck posed by autoregressive rollout generation in post-training reinforcement learning. It systematically integrates state-of-the-art speculative decoding techniques—including Eagle3, pretrained multi-token prediction (MTP) heads, and compact draft models—into the RL training pipeline using the NeMo-RL framework with a vLLM inference backend. The approach enables both synchronous and asynchronous rollout inference while preserving the target model’s output distribution. Experimental results demonstrate a 1.8× increase in rollout throughput for synchronous training with an 8B-parameter model. When combined with asynchronous mechanisms, the method is projected to accelerate end-to-end training by up to 2.5× at the 235B-parameter scale.
This study addresses the prohibitive decoding costs associated with long chain-of-thought reasoning. We propose Trajectory Locally Adaptive Retrieval (TLAR), a method that repurposes a model's historical trajectories as runtime memory to accelerate generation. TLAR introduces a novel adaptive retrieval mechanism that dynamically adjusts the activation threshold and candidate tree width, coupled with an exact verification algorithm. This combination optimizes speculative decoding while strictly preserving output distribution consistency. Experimental evaluations demonstrate that TLAR substantially improves token acceptance rates and end-to-end throughput across code debugging, mathematical reasoning, and writing tasks, achieving a significant breakthrough in inference efficiency.
本文提出选择性再生解码方法,通过保留部分候选轨迹的高质量前缀并仅修正退化的后缀,提高推理时的计算效率和轨迹质量。
研究通过Rollout-Decoded Reconstruction方法,在混沌系统中提高了长期预测的准确性,通过在训练时引入自由运行模型并惩罚重建误差来解决预测差距问题。
This work addresses the inefficiency of ODE-based samplers in diffusion and flow-matching models, which typically require tens to hundreds of neural network evaluations per sample. While existing acceleration methods often rely on retraining or distillation, the authors propose Truncated Jump Sampling (TJS), a training-free early-exit strategy. Building on the x-prediction perspective and leveraging the information exposure of intermediate states about the initial sample \(x_0\) along affine probability paths, they formalize the novel concept of “endpoint decodability” and prove its equivalence to minimum mean squared error estimation. TJS exploits this insight to significantly accelerate sampling without altering model architecture or requiring retraining. Experiments demonstrate that TJS reduces the number of function evaluations (NFE) by 20%–70% across SDXL, SD3.5M, Z-Image-Turbo, and multiple class-conditional benchmarks while preserving near-original generation quality.
This study addresses the throughput-latency conflicts and dynamic reallocation dependencies inherent in scheduling long trajectories for reinforcement learning (RL) post-training across heterogeneous accelerators. To this end, we propose CadenceRL, a framework that achieves hardware-specialized scheduling through structural workload reshaping. By substituting per-move evaluation with cadence control and centralized strategies, CadenceRL effectively breaks scheduling circular dependencies. Furthermore, it integrates asynchronous RL and KV-cache pre-allocation to automatically align heterogeneous resources with their optimal roles. Experimental results demonstrate that, without manual routing configurations, CadenceRL improves decoding throughput by 48% and reduces P95 trajectory latency by 64%.
研究通过对比不同预测模型在闭环控制下的表现,提出评估机器人预测模型应考虑预测时长和测量更新频率。
该研究针对多轮代理轨迹中的冗余问题,通过构建依赖DAG来精简轨迹,并基于此进行微调,以提高效率和准确性。
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
This study addresses the reliability-cost imbalance in terminal agent training caused by undefined supervision horizon lengths, proposing a selective long-horizon refinement method. Grounded in a bias-complexity theoretical analysis that reveals performance saturation, this work establishes a new paradigm demonstrating trajectory selection's superiority over full-scale training. A two-stage progressive strategy is designed to first warm-start the model with short prefixes and subsequently perform selective continuation optimization via probabilistic filtering. This approach significantly enhances task success rates and stability while reducing training overhead by 30%, substantially outperforming existing baselines across multiple benchmarks.