🤖 AI Summary
Autoregressive decoding in large language models (LLMs) suffers from inherent sequential bottlenecks and high latency for long outputs. Method: This paper proposes a model-intrinsic, training-free parallel decoding architecture based on a novel “speculative consensus coordination” paradigm: multiple generation streams collaborate via a shared dynamic latent space and broadcasted semantic “notes”; it introduces a lightweight Speculative Note Conditioning (SNC) adapter, a learnable verification head, and a global semantic bus, trained progressively over 50k steps. Contribution/Results: On a frozen 20B-parameter model, the method achieves 77.8% coverage prediction accuracy and near-serial semantic recovery fidelity—without any weight modification. Unlike external orchestration approaches (e.g., Skeleton-of-Thought), it eliminates coherence drift caused by inter-stream communication deficits, significantly improving efficiency for long-text generation.
📝 Abstract
Autoregressive decoding in Large Language Models (LLMs) is inherently sequential, creating a latency bottleneck that scales linearly with output length. While ``Decomposition-and-Fill'' methods like Skeleton-of-Thought attempt to parallelize generation via external orchestration, they suffer from extit{coherence drift} due to the lack of cross-stream communication. In this work, we introduce the extbf{Parallel Decoder Transformer (PDT)}, a parameter-efficient architecture that embeds coordination primitives directly into the inference process of a frozen pre-trained model.
Instead of retraining the base model, PDT injects lightweight extit{Speculative Note Conditioning (SNC)} adapters that allow parallel decoding streams to synchronize via a shared, dynamic latent space. We formulate coordination as a extit{speculative consensus} problem, where sibling streams broadcast semantic ``notes'' to a global bus, gated by a learned verification head. We validate our approach on a 50,000-step curriculum using a frozen 20B-parameter backbone. Our results demonstrate that PDT achieves effective self-correction, reaching extbf{77.8% precision} in coverage prediction and recovering approximate serial semantics without modifying the trunk weights. This establishes PDT as a scalable, efficient alternative to full model fine-tuning for structured parallel generation.