confidence-gated decoding

Designs, builds, or analyzes two-stage decoding pipelines that use a lightweight feed‑forward neural decoder as a fast path to produce predictions and per‑sample confidence estimates, and a heavier refinement stage that is invoked only when the confidence falls below a threshold or ordering policy. Work includes defining confidence estimators and thresholds/ordering rules, trading off accuracy versus latency and throughput, minimizing costly refinement invocations, and implementing CPU‑friendly, low‑latency neural decoders.

confidence-gateddecoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding

May 02, 2025
BM
Bradley McDanel
🏛️ Franklin and Marshall College | New York University | University of Pennsylvania

Low hardware utilization and throughput bottlenecks plague large language model (LLM) inference due to sequential stage dependencies. To address this, we propose a multi-level hierarchical pipelined speculative decoding framework that enables asynchronous collaboration—prediction, verification, and dynamic rollback—among *k* heterogeneous models, overcoming the temporal constraints of conventional single-stage speculation. We introduce the first hierarchical pipeline architecture, derive a theoretical lower bound on throughput, and present the first closed-form analytical solution for steady-state verification probability, uncovering the intrinsic “depth–gain” trade-off. Leveraging a lightweight coordination mechanism and adaptive rollback strategy, our method achieves up to 2.54× end-to-end speedup on LLaMA-2/3, significantly outperforming state-of-the-art approaches. Extensive evaluation on text summarization and code generation demonstrates cross-task and multi-device scalability, with system efficiency monotonically improving as pipeline depth increases.

Enhancing hardware utilization via hierarchical pipeline modelsOvercoming sequential stage dependencies in speculative decodingScalably accelerating LLM inference on multi-device systems

Existing speculative decoding methods rely on serial multi-token prediction, which leads to progressively increasing prediction difficulty and pipeline bubbles, thereby limiting acceleration in low-concurrency scenarios. This work proposes a speculative pipeline decoding framework that deeply integrates model pipeline parallelism into speculative decoding for the first time. By partitioning a large language model into multiple stages that process tokens in parallel, the framework introduces a parallel speculative module strictly synchronized with the main model and incorporates cross-stage intermediate feature aggregation alongside an efficient verification mechanism. This approach achieves bounded prediction complexity, high token acceptance rates, and zero bubble latency, significantly improving theoretical speedup and offering a highly scalable and efficient decoding solution for large language model inference.

decoding accelerationLLM inferencemulti-token prediction

Decoding Speculative Decoding

Feb 02, 2024
MY
Minghao Yan
🏛️ University of Wisconsin-Madison

This work addresses the critical challenge of draft model selection in speculative decoding, revealing that inference latency—not language modeling capability—is the primary bottleneck limiting acceleration. We therefore propose a hardware-aware draft model design paradigm: abandoning conventional language-modeling–centric selection criteria in favor of optimizing explicitly for low latency and high throughput via lightweight architecture customization and co-optimization with the inference engine. Evaluated across over 350 configurations on LLaMA-65B and OPT-66B, our approach achieves a 111% throughput improvement for draft models, substantially outperforming existing methods. Furthermore, we demonstrate strong generalization across diverse foundation models—including LLaMA-1, LLaMA-2, LLaMA-3.1—and multiple supervised fine-tuned variants. Our framework establishes a new pathway toward efficient, general-purpose, and hardware-adapted large language model inference.

Analyzing factors affecting speculative decoding performanceDesigning hardware-efficient draft models for LLMsOptimizing draft model selection for speculative decoding

Boosting Ordered Statistics Decoding of Short LDPC Codes With Simple Neural Network Models

Apr 22, 2024
GL
Guang-He Li
🏛️ Shandong Technology and Business University | Binzhou Medical University

Short LDPC codes exhibit residual errors under normalized min-sum (NMS) decoding, while conventional ordered statistics decoding (OSD) achieves high error-correction performance at prohibitively high computational complexity—rendering it unsuitable for ultra-low-latency applications. To address this trade-off, we propose a lightweight neural-enhanced OSD framework. Our method introduces: (i) a sliding-window-assisted neural network to enable early termination of OSD iterations; (ii) a bit-reliability refinement mechanism leveraging iterative failure information; and (iii) a structured, block-wise test-error pattern generation strategy. Evaluated on standardized short LDPC codes, the proposed decoder attains state-of-the-art (SOTA) bit-error-rate (BER) performance—matching or surpassing conventional OSD—while reducing average decoding latency by up to 62% and computational complexity by over 50%. The approach is particularly well-suited for ultra-reliable low-latency communication (URLLC) systems, such as those in 5G and beyond.

Balancing latency and accuracy with neural networksImproving decoding performance for short LDPC codesReducing computational complexity in ordered statistics decoding

Latest Papers

What's happening recently
View more

This work addresses the efficiency bottleneck in speculative decoding caused by distributional mismatch between general-purpose draft models and downstream tasks. To overcome this limitation, the authors propose a task-adaptive, lightweight training and fusion strategy for draft models. Specifically, they fine-tune HASS and EAGLE-2 architectures on task-specific datasets such as MathInstruct and ShareGPT, and integrate a confidence-based routing mechanism with a merged-tree verification approach. This design significantly improves both the acceptance length of speculative tokens and task adaptability. Experimental results demonstrate that task-specialized draft models achieve superior performance on their respective benchmarks, while mixed-task training enhances robustness. Moreover, the proposed fusion strategy outperforms conventional weight averaging, delivering state-of-the-art overall results across multiple evaluation metrics.

draft modelinference-time combinationspeculative decoding

This work addresses the latency and throughput bottlenecks in real-time quantum error correction decoding during the transition from NISQ to fault-tolerant quantum computing. It proposes an adaptive confidence-gated two-stage decoding framework, wherein a lightweight feedforward neural network rapidly processes high-confidence syndromes, while only 3.3%–6.2% of low-confidence samples are refined using minimum-weight perfect matching (MWPM). This approach introduces confidence gating for the first time in rotated surface code decoding, achieving a logical accuracy improvement from 99.21% to 99.81% across code distances d = 3–11, while delivering a high throughput of 4.6×10⁵ samples per second on CPU. The study further establishes a hardware-aware co-design paradigm that jointly optimizes decoding accuracy, computational efficiency, and deployability.

hardware-awarelatency-constrainedquantum error correction

This study addresses the inference inefficiency of large language models caused by memory bandwidth constraints and limited parallelism in single-stream autoregressive decoding. Conducting the first device-agnostic empirical investigation of speculative decoding on consumer-grade Apple Silicon hardware, the work proposes an acceleration framework that leverages a small draft model to pre-generate multiple tokens, followed by batch verification and rejection sampling using the target model. The approach strictly preserves output distribution equivalence, as validated by chi-squared tests and sequence consistency checks. Experimental results demonstrate that the optimal configuration (K=6) achieves a 1.61× speedup with a 37.8% acceptance rate, while also revealing three configurations that incur slowdowns due to pseudo-parallelization overhead or insufficient draft model quality.

autoregressive decodingconsumer hardwarelarge language models

This work addresses the fundamental trade-off between microsecond-scale latency and high decoding accuracy that hinders the practical deployment of neural decoders in quantum error correction. Under explicit accuracy–latency constraints, the authors unify and reconstruct five representative surface code neural decoder architectures, proposing an end-to-end compression pipeline and demonstrating FPGA deployment supporting code distances up to $d=9$. Their findings reveal that recent gains in decoding performance are primarily driven by training data volume rather than architectural complexity, that effective inductive bias is crucial for achieving high accuracy, and that INT4 quantization is essential to meet microsecond-level latency requirements. This study provides a viable technical pathway and empirical foundation for scalable, real-time neural quantum error correction.

accuracy-latency tradeoffneural decodersquantum error correction

This study systematically evaluates the practical performance advantages of emerging AI accelerators over GPUs in large language model (LLM) inference by decoupling the process into Prefill and Decode phases—a first in the field. Using the Llama2-7B model, the authors conduct phase-specific benchmarking and heterogeneous deployment experiments on platforms including GPU and GroqRack, measuring Time-To-First-Token (TTFT) and Time-Per-Output-Token (TPOT). Results show that GPUs consistently outperform in the Prefill phase, while GroqRack achieves lower TPOT in the Decode phase under no-batching conditions; however, GPUs surpass GroqRack in throughput under high batching. The work delineates the effective conditions for heterogeneous task partitioning, offering actionable insights for hardware selection and scheduling in LLM inference systems.

AI acceleratorslatency metricsLLM inference

Hot Scholars

YF

Yanwei Fu

Fudan University
Computer visionmachine learningMultimedia
CZ

Changqing Zhang

Professor, Tianjin University
Machine LearningMultimodal LearningLLM
ZL

Ziquan Liu

Assistant Professor, Queen Mary University of London
machine learning
AC

Andrzej Cichocki

Systems Research Institute, Nicolaus Copernicus University, RIKEN (AIP)
Biomedical Signal ProcessingNeural EngineeringTensor DecompositionAnalog_Electronics
WW

Wenwen Wang

Assistant Professor, School of Computing, University of Georgia
Computer Systems