Parallel Decoder Transformer: Model-Internal Parallel Decoding with Speculative Invariance via Note Conditioning

📅 2025-12-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Autoregressive decoding in large language models (LLMs) suffers from inherent sequential bottlenecks and high latency for long outputs. Method: This paper proposes a model-intrinsic, training-free parallel decoding architecture based on a novel “speculative consensus coordination” paradigm: multiple generation streams collaborate via a shared dynamic latent space and broadcasted semantic “notes”; it introduces a lightweight Speculative Note Conditioning (SNC) adapter, a learnable verification head, and a global semantic bus, trained progressively over 50k steps. Contribution/Results: On a frozen 20B-parameter model, the method achieves 77.8% coverage prediction accuracy and near-serial semantic recovery fidelity—without any weight modification. Unlike external orchestration approaches (e.g., Skeleton-of-Thought), it eliminates coherence drift caused by inter-stream communication deficits, significantly improving efficiency for long-text generation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Autoregressive decoding in Large Language Models (LLMs) is inherently sequential, creating a latency bottleneck that scales linearly with output length. While ``Decomposition-and-Fill'' methods like Skeleton-of-Thought attempt to parallelize generation via external orchestration, they suffer from extit{coherence drift} due to the lack of cross-stream communication. In this work, we introduce the extbf{Parallel Decoder Transformer (PDT)}, a parameter-efficient architecture that embeds coordination primitives directly into the inference process of a frozen pre-trained model. Instead of retraining the base model, PDT injects lightweight extit{Speculative Note Conditioning (SNC)} adapters that allow parallel decoding streams to synchronize via a shared, dynamic latent space. We formulate coordination as a extit{speculative consensus} problem, where sibling streams broadcast semantic ``notes'' to a global bus, gated by a learned verification head. We validate our approach on a 50,000-step curriculum using a frozen 20B-parameter backbone. Our results demonstrate that PDT achieves effective self-correction, reaching extbf{77.8% precision} in coverage prediction and recovering approximate serial semantics without modifying the trunk weights. This establishes PDT as a scalable, efficient alternative to full model fine-tuning for structured parallel generation.
Problem

Research questions and friction points this paper is trying to address.

Parallelizes autoregressive decoding to reduce latency in LLMs.
Prevents coherence drift in parallel generation streams via internal coordination.
Enables structured parallel decoding without retraining the base model.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Decoder Transformer enables internal parallel decoding
Speculative Note Conditioning adapters synchronize streams via latent space
Learned verification head gates semantic notes for speculative consensus
🔎 Similar Papers
No similar papers found.
Independent Researcher
L
Logan Robbins
Independent Researcher