Score
Design and build decoder architectures and associated decoding algorithms that transform encoded representations into target outputs, covering both learned and classical decoders such as belief‑propagation, dynamic‑programming, convolutional, and encoder–decoder formulations. This work includes implementing iterative, parallel, non‑autoregressive, sliding‑window and speculative decoding mechanisms, enforcing constraint‑aware rules, integrating decoders with encoders or other network outputs, and analyzing tradeoffs in error floor, latency, computational and memory efficiency for lightweight or hardware‑friendly deployment.
Neural decoding of LDPC codes faces challenges in modeling long-range dependencies over graph structures and effectively incorporating the principles of classical belief propagation (BP). Method: This paper proposes a novel Transformer-based neural decoder. It introduces a differentiable syndrome loss to enforce parity-check constraints end-to-end, and a differential attention mechanism that explicitly models bidirectional message passing between bit and check nodes—embedding BP’s message update rules into self-attention computations. Additionally, graph-structure-aware node embeddings and hierarchical supervision signals enable joint learning of the LDPC code’s global topology. Results: Experiments demonstrate that the proposed decoder significantly outperforms conventional BP and state-of-the-art neural decoders on short-to-medium-length LDPC codes, achieving 0.5–1.2 dB coding gain in bit error rate under AWGN channels.
To address the degraded decoding performance of linear block codes under additive white Gaussian noise and nonlinear optical fiber channels, this paper proposes an adaptive weighted belief propagation (WBP) decoding algorithm that overcomes the limitations of fixed edge weights by dynamically optimizing weights for each received codeword. Two lightweight adaptive architectures are innovatively introduced: (1) a parallel discrete weight search mechanism leveraging an unrolled Tanner graph with trainable edge weights; and (2) a compact neural network embedded as a real-time weight generator enabling online adaptation. Extensive evaluations across diverse linear block codes demonstrate a one-order-of-magnitude reduction in bit error rate. In long-haul optical communication scenarios, the proposed method achieves 0.8 dB coding gain over the neural normalized min-sum algorithm while maintaining comparable computational complexity.
This study addresses the inference inefficiency of large language models caused by memory bandwidth constraints and limited parallelism in single-stream autoregressive decoding. Conducting the first device-agnostic empirical investigation of speculative decoding on consumer-grade Apple Silicon hardware, the work proposes an acceleration framework that leverages a small draft model to pre-generate multiple tokens, followed by batch verification and rejection sampling using the target model. The approach strictly preserves output distribution equivalence, as validated by chi-squared tests and sequence consistency checks. Experimental results demonstrate that the optimal configuration (K=6) achieves a 1.61× speedup with a 37.8% acceptance rate, while also revealing three configurations that incur slowdowns due to pseudo-parallelization overhead or insufficient draft model quality.
This work addresses the fundamental limitation of single- and multi-layer neural network (SLNN/MLNN) decoders—namely, their inability to guarantee optimal performance and reliance on approximate, empirically trained solutions. We propose an analytical, training-free decoder with strictly binary weights, directly constructed from the linear block code’s parity-check or generator matrix as a deterministic combinational logic circuit. Crucially, this circuit is provably equivalent to maximum-likelihood (ML) decoding. We provide the first theoretical proof that both single- and multi-label architectures achieve strict ML performance. Empirical evaluation on short codes—including Hamming(7,4), Polar(16,8), and BCH(31,21)—confirms attainment of the theoretical minimum bit-error rate, with significantly lower computational complexity than state-of-the-art neural decoders. Our core contribution lies in eliminating parameter learning and mitigating the curse of dimensionality by establishing a verifiably optimal, interpretable mapping from codebook to binary circuit—yielding the first neuromorphic decoding paradigm for short codes that is simultaneously explainable, deterministic, and ML-optimal.
To address the high inference latency induced by autoregressive decoding in large language models (LLMs), this paper proposes a novel speculative decoding (SD) paradigm. The method introduces a lightweight draft model coupled with multi-token parallel sampling, enabling rapid draft generation and concurrent verification via a dedicated validation module; it further incorporates probabilistic consistency calibration to preserve output distribution fidelity. Crucially, this work achieves the first tight integration of draft generation and parallel verification—enabling 2–4× end-to-end speedup without compromising generation quality. A systematic analysis explores the SD architectural design space, validates the efficacy of verification strategies, and characterizes scalability limits. The approach is plug-and-play, fully compatible with both open-source and industrial-grade LLM deployments. By bridging theoretical insight with practical implementation, it delivers a production-ready pathway for efficient LLM inference.
This study addresses the lack of systematic evaluation of energy consumption characteristics in speculative decoding strategies for large language models. It presents the first comprehensive quantification of fine-grained energy usage across diverse speculative decoding methods, examining their performance under varying model scales and architectures, decoding strategies, and datasets. The work uncovers the synergistic effects of model design, algorithmic choices, and data properties on energy efficiency, identifying key factors that determine the energy efficacy of speculative decoding. These findings provide empirical grounding and actionable insights for optimizing large-model inference toward lower energy consumption.
This work addresses the high decoding complexity and suboptimal performance of conventional belief propagation (BP) decoding for BCH codes compared to LDPC codes. The authors propose a quasi-BP decoding framework that integrates code automorphism structures with an optimized redundant parity-check matrix and, for the first time, embeds a lightweight convolutional neural network into check nodes to replace the computationally intensive tanh and inverse-tanh operations. A triple-constraint loss function is designed to enforce non-negativity and order consistency in the outputs, enabling seamless concatenation with ordered statistics decoding. Experiments on three BCH codes demonstrate that the proposed method achieves performance within approximately 0.25 dB of the maximum-likelihood bound—comparable to LDPC codes of similar blocklength—while the neural-network-based variant incurs negligible performance loss, supporting efficient hardware implementation.
This work addresses the limitations of existing speculative decoding methods in long-context scenarios, where static heuristics fail to adapt to the dynamic computational overhead of attention mechanisms. We formalize draft model selection as a knapsack optimization problem and propose a novel framework that decouples attention and MLP layers, constructs a context-aware hardware latency model, and employs a parallel dynamic programming algorithm to select optimal draft configurations in real time for maximal throughput. To ensure draft fidelity without retraining, we introduce a theoretical guarantee based on cosine similarity and develop a training-free, adaptive layer selection mechanism. Evaluated on Qwen3 and Llama3, our approach achieves up to 1.47× end-to-end speedup over state-of-the-art methods while preserving the target model’s output distribution.
Autoregressive decoding in large language models (LLMs) suffers from inherent sequential bottlenecks and high latency for long outputs. Method: This paper proposes a model-intrinsic, training-free parallel decoding architecture based on a novel “speculative consensus coordination” paradigm: multiple generation streams collaborate via a shared dynamic latent space and broadcasted semantic “notes”; it introduces a lightweight Speculative Note Conditioning (SNC) adapter, a learnable verification head, and a global semantic bus, trained progressively over 50k steps. Contribution/Results: On a frozen 20B-parameter model, the method achieves 77.8% coverage prediction accuracy and near-serial semantic recovery fidelity—without any weight modification. Unlike external orchestration approaches (e.g., Skeleton-of-Thought), it eliminates coherence drift caused by inter-stream communication deficits, significantly improving efficiency for long-text generation.
To address high inference latency, quantization-induced accuracy degradation, and non-negligible overhead from existing speculative decoding methods for large language models (LLMs), this paper proposes SPEQ—a training-free, zero-storage-overhead, lossless speculative decoding framework. SPEQ’s core innovations are: (1) dynamic lightweight draft model construction via floating-point exponent remapping and parameter sharing, reusing original model weights without additional parameters; and (2) a novel algorithm-hardware co-design enabling unified execution of draft generation and verification on reconfigurable computing arrays. SPEQ incurs no training cost or storage overhead while achieving both high throughput and zero accuracy loss. Extensive experiments across 15 LLMs and diverse tasks demonstrate that SPEQ accelerates inference by 2.07×, 1.53×, and 1.45× over FP16 baseline, Olive, and Tender, respectively.