Score
Designs, implements, and evaluates models and training/inference pipelines that factorize joint sequence distributions into sequential conditionals—this includes decoder-only transformer autoregressive models, discrete and continuous (vector) autoregression, hierarchical, block‑autoregressive, semi‑autoregressive, and hybrid parallel‑autoregressive architectures and their sampling/decoding procedures. Also builds the associated loss functions and supervision methods (e.g., next‑token loss, LLM‑supervised retrieval and embedding alignment without contrastive pairs), defines token‑ordering schemes (e.g., base‑first or base/physical decoding), and analyzes how these choices affect multi‑step unrolled training, conditional generation, and retriever integration.
This work addresses the limitations of traditional autoregressive models, which rely solely on next-token prediction and struggle to capture sequence-level properties, often resulting in local overfitting and poor global structure modeling. Moreover, controllable generation typically demands costly sampling or architectural modifications. To overcome these challenges, the authors propose the Conditional Attribute Transformer, which jointly optimizes next-token prediction and sequence attribute estimation conditioned on the current token within a single forward pass. This approach uniquely enables token-level attribute attribution, counterfactual attribute evaluation, and attribute-guided generation—all within a unified model without additional sampling or structural changes. Experiments demonstrate state-of-the-art performance on sparse-reward tasks, improved language modeling at large scales, attribute estimation orders of magnitude faster than sampling-based methods, and effective guidance across diverse language generation tasks.
Autoregressive decoding in large language models incurs high latency, and existing multi-token prediction methods rely on strong independence assumptions that limit modeling fidelity. Method: We propose Parallel Token Prediction (PTP), the first framework to internalize the sampling process into the model architecture, enabling joint generation of multiple semantically coherent tokens in a single Transformer forward pass while strictly preserving expressivity over any autoregressive distribution—thereby eliminating restrictive independence assumptions. PTP integrates inverse autoregressive training with explicit sampling modeling and supports both teacher-free and distillation-based training. Results: Evaluated on Vicuna-7B, PTP achieves 4.12 average accepted tokens per speculative step on Spec-Bench, maintains full modeling capability for long-sequence generation, and attains state-of-the-art performance in speculative decoding.
Autoregressive normalizing flows suffer from slow inference due to strict sequential dependencies that preclude parallelization. This paper proposes Selective Jacobian Decoding (SeJD), the first method to empirically identify and exploit a hierarchical redundancy pattern in autoregressive generation: lower-layer dependencies exhibit low redundancy, whereas higher-layer ones are highly redundant—enabling relaxation of full-sequence dependency constraints. SeJD integrates local Jacobian approximation, hierarchical dependency modeling, and parallel iterative optimization, yielding a decoding mechanism with theoretically guaranteed superlinear convergence and iteration count bounded by that of sequential sampling. Experiments across multiple datasets demonstrate up to 4.7× inference speedup without compromising generation quality or fidelity. The core contributions are: (i) uncovering the layered dependency redundancy property in autoregressive flows; and (ii) introducing the first parallel autoregressive flow decoding framework that jointly ensures rigorous convergence guarantees and practical acceleration.
Autoregressive vision-language models suffer from high inference latency—scaling linearly with sequence length O(L)—hindering practical deployment. Method: This paper proposes the first non-autoregressive sequence-to-sequence framework for image-to-text generation. Its core innovation is the introduction of Query-CTC loss into vision-language modeling, enabling direct optimization of the joint token distribution (rather than conditional distributions) and achieving constant-time O(1) parallel decoding. The approach integrates latent variable path marginalization with end-to-end training, eliminating the need for predefined alignments or external aligners. Results: Experiments demonstrate that the method matches the accuracy of state-of-the-art autoregressive models across diverse vision-language understanding and generation tasks, while significantly accelerating inference. It thus achieves a favorable trade-off between efficiency and performance without compromising fidelity.
Autoregressive models suffer from high real-time inference latency due to sequential dependency, and conventional compression techniques—such as pruning and quantization—often incur significant accuracy degradation. To address this, we propose a unified generation–refinement decoding framework. Methodologically, we first establish a taxonomy of generation strategies—including n-gram matching and draft-model-based approaches—and refinement mechanisms—spanning single-step verification and iterative optimization. The framework integrates speculative decoding, multi-round verification, knowledge distillation from draft models, and hardware-aware scheduling for efficient deployment across heterogeneous platforms. Evaluated on text, image, and speech generation tasks, our approach achieves an average 2.1× speedup in end-to-end latency while sustaining minimal accuracy loss (<0.5% in BLEU, CLIP, and FID metrics). This work provides both scalable theoretical foundations and system-level implementation strategies for real-time large language and multimodal model applications.
This work addresses the limitations of traditional autoregressive models, which rely on discrete tokenization and struggle to accurately model continuous values—often leading to functional failures in precision-sensitive tasks such as semiconductor circuit design. To overcome this, the authors propose AGDC, a unified framework that enables end-to-end autoregressive generation of hybrid discrete-continuous sequences for the first time. Built upon a Transformer architecture, AGDC integrates classification-based prediction with diffusion modeling and introduces a dynamic EOS logit adjustment mechanism alongside a length regularization loss. Evaluated on a newly curated high-precision semiconductor layout dataset, ContLayNet (334K samples), and SVG-based graphics tasks, AGDC significantly outperforms both discretization-based and fixed-structure baselines, breaking through the precision bottleneck and enabling high-fidelity, variable-length vector data generation.
This work addresses a key limitation of existing semi-autoregressive draft models—such as DSpark—which employ linear generation structures wherein early token mismatches cause complete failure of subsequent predictions, thereby constraining the acceleration potential of large draft blocks. To overcome this, the authors propose a tree-based draft generation method that leverages a pretrained Markov head to independently score multiple child tokens for each parent node and prioritizes verification of higher-probability paths under a fixed verification budget. Notably, this approach requires no retraining or additional backbone forward passes, achieving improved acceleration solely through inference-time modifications. Evaluated on the Qwen3 model series across nine benchmarks, the method yields relative speedup gains of 3.1%–29.5% over DSpark; specifically, on GSM8K, Qwen3-4B increases its average accepted length from 9.41 to 11.16, raising the speedup ratio from 6.14× to 6.60×.
Autoregressive decoding in large language models (LLMs) suffers from inherent sequential bottlenecks and high latency for long outputs. Method: This paper proposes a model-intrinsic, training-free parallel decoding architecture based on a novel “speculative consensus coordination” paradigm: multiple generation streams collaborate via a shared dynamic latent space and broadcasted semantic “notes”; it introduces a lightweight Speculative Note Conditioning (SNC) adapter, a learnable verification head, and a global semantic bus, trained progressively over 50k steps. Contribution/Results: On a frozen 20B-parameter model, the method achieves 77.8% coverage prediction accuracy and near-serial semantic recovery fidelity—without any weight modification. Unlike external orchestration approaches (e.g., Skeleton-of-Thought), it eliminates coherence drift caused by inter-stream communication deficits, significantly improving efficiency for long-text generation.
This work addresses the absence of a unified theoretical framework for autoregressive decoding strategies in speech processing, which has led to ambiguous definitions, inconsistent taxonomies, and difficulties in fair comparison. The paper introduces, for the first time, a general formal framework that precisely specifies inclusion criteria for autoregressive search and systematically categorizes and describes decoding strategies employed in neural speech generation models. By clarifying conceptual boundaries, the framework enhances comparability and evaluation consistency across strategies, streamlines the design of decoding-centric benchmarking protocols, and enables ablation studies focused specifically on search mechanisms. Consequently, it facilitates standardized analysis of inference-stage behavior in speech generation models.
This study investigates the fundamental limits of language models in learning the true data-generating process of natural language solely from observed text sequences, particularly when unobserved contextual factors—such as facts, intentions, or social settings—influence generation. By distinguishing between the full conditional generative process, the marginalized text-only process, and the distribution learned by the model, the work introduces criteria based on local sufficient statistics and conditional mutual information to delineate when next-token prediction remains valid. It demonstrates that standard training implicitly assumes stationarity and ergodicity, assumptions often violated in heterogeneous corpora, and proves that marginal models succeed only when observed prefixes are approximately sufficient for latent variables. The paper further interprets retrieval-augmented generation (RAG) and tool use as mechanisms that restore conditional sufficiency, thereby transcending conventional language modeling paradigms.