Score
Designs, builds, and analyzes decoder architectures and decoding-algorithm implementations — including probabilistic decoders — that minimize computation, memory, and hardware resource use. Work focuses on concrete engineering choices to reuse computation and memory over time, limit hardware growth with scaling, and reduce memory conflicts and pipeline stalls via multi-bank hashing, hierarchical ID mapping, and related techniques.
In fault-tolerant quantum computing (FTQC), classical decoders for quantum error correction (QEC) face severe resource demand fluctuations—peak loads can exceed idle-period requirements by several orders of magnitude—rendering static hardware allocation inefficient (causing underutilization or real-time violations). Method: We propose the first workload-aware decoder virtualization framework and latency-aware dynamic scheduling strategy, enabling elastic, on-demand allocation of hardware decoder resources. Our approach integrates fine-grained syndrome processing modeling, reusable customized decoder designs, and cross-logical-qubit task coordination. Contribution/Results: Evaluated at the 100–1,000 logical qubit scale, our method reduces hardware decoder resource requirements by 10× while strictly guaranteeing real-time decoding latency and fault-tolerance reliability. This breakthrough overcomes a critical scalability bottleneck in the classical processing layer of QEC, providing essential infrastructure for practical FTQC deployment.
This work addresses the memory-bandwidth-bound nature of large language model (LLM) decoding, which leads to low GPU utilization and high inference costs. The authors propose Skymizer HTX-301, a dedicated decoding accelerator that, for the first time, quantifies at the chip level the two primary decoding bottlenecks—fill-to-batch (F/B) and fill-to-sequence (F/S). By deliberately reducing computational density and leveraging cost-effective DDR5 memory and PCIe interfaces within a 28nm process node, the design achieves high cost efficiency without relying on advanced semiconductor nodes or high-bandwidth memory. A single HTX-301 card, priced at approximately $19,000, supports deployment of a 671B-parameter model and delivers stable dual-user throughput of 20.3 tokens per second in a 4U server configuration. This approach reduces inference costs to one-twelfth that of an H100-based solution, amounting to just $12 per million tokens.
To address high latency and energy consumption in large language model (LLM) inference caused by increasing model complexity, this paper proposes a hardware-native speculative decoding architecture. Unlike conventional software-level speculation, our approach deeply integrates speculative execution into the hardware accelerator design, enabling instruction-level speculative execution and overcoming synchronization and latency bottlenecks. The architecture comprises a RISC-V-based custom LLM acceleration core, a dedicated hardware predictor, a lightweight draft model co-scheduling circuit, and a dynamic cache coherence protocol. Evaluated on Llama-2-7B, the design achieves a 2.8× end-to-end throughput improvement, a 3.4× energy efficiency gain, and a 41% reduction in first-token latency—significantly outperforming state-of-the-art baselines including vLLM and SpecInfer.
High-level synthesis (HLS) tools—such as Chisel and commercial HLS compilers—often produce circuits with inferior performance compared to hand-designed hardware in high-performance computing (HPC) accelerator design. Method: This paper proposes a hierarchical algorithmic decomposition and automated evaluation framework that abstracts mathematical kernels (e.g., Fourier transforms, matrix multiplication, QR decomposition) into reusable building blocks, uniformly implemented across multiple abstraction levels (RTL, Chisel, C++ HLS), and systematically benchmarked for resource utilization, timing, and operating frequency. Contribution/Results: The framework establishes the first cross-abstraction-level fair benchmarking methodology, enabling fine-grained identification of inefficiencies introduced by HLS compilers. Experimental evaluation demonstrates significantly improved accuracy and interpretability in pinpointing design bottlenecks, providing quantitative guidance for both HLS tool optimization and practical hardware design.
This work addresses the challenge of automatically mapping and efficiently solving combinatorial optimization problems on probabilistic computing architectures. It proposes an Ising-model-based automated mapping approach that translates optimization problems into hardware-compatible formulations by constructing corresponding Hamiltonians and configuring the number of p-bits. A key innovation is the introduction of an adaptive algorithm selection mechanism that dynamically switches among Gibbs sampling, simulated annealing, simulated quantum annealing, and cluster updates to significantly enhance convergence speed and robustness. Validated through MTJ-based hardware modeling on standard benchmark instances, the proposed framework demonstrates superior performance over fixed-strategy approaches, offering a scalable and systematic co-design and evaluation paradigm for probabilistic computing systems.
This work addresses the challenge of achieving high-throughput, low-latency hardware implementations of BCH decoders that simultaneously support flexible error-correction capabilities and arbitrary code lengths. The authors propose two novel architectures: a conventional decoder based on the Berlekamp–Massey algorithm and Chien search, which accommodates any error-correction capability and code length, and a pioneering direct decoder that efficiently computes the roots of the error-locator polynomial, supporting up to t = 4 errors. Both architectures are optimized for Xilinx Ultrascale+ FPGAs and 16 nm FinFET technology. In 16 nm implementation, the (256,239) and (256,223) codes achieve single-cycle decoding with throughputs of 239 Gb/s and 223 Gb/s, respectively, and latencies of only 2–8 ns, substantially improving area efficiency and throughput performance.
This work addresses the lack of systematic understanding in co-designing the prefill and decode phases for heterogeneous large language model inference, which hinders deployment efficiency. Focusing on four key design axes—accelerator architecture, precision, interconnect, and KV residency—the study reveals that only a subset of these factors forms strongly coupled constraints and identifies three pivotal boundary decisions: compute placement, KV representation, and KV ownership. Integrating empirical insights from industrial deployments with runtime analysis, the authors propose a runtime role–driven precision strategy, a byte-level KV transfer mechanism, and explicit ownership management to construct a multi-axis design space model. The resulting framework yields actionable design guidelines that substantially improve heterogeneous inference efficiency.
This study addresses the inference inefficiency of large language models caused by memory bandwidth constraints and limited parallelism in single-stream autoregressive decoding. Conducting the first device-agnostic empirical investigation of speculative decoding on consumer-grade Apple Silicon hardware, the work proposes an acceleration framework that leverages a small draft model to pre-generate multiple tokens, followed by batch verification and rejection sampling using the target model. The approach strictly preserves output distribution equivalence, as validated by chi-squared tests and sequence consistency checks. Experimental results demonstrate that the optimal configuration (K=6) achieves a 1.61× speedup with a 37.8% acceptance rate, while also revealing three configurations that incur slowdowns due to pseudo-parallelization overhead or insufficient draft model quality.
This work addresses the limitations of traditional hardware prefetchers, which rely on manual analysis of program traces and struggle to handle performance anomalies in real-world workloads. The authors propose a performance-anomaly-driven automated design methodology that employs an intelligent agent to diagnose prefetching failures, automatically synthesize a Mixture-of-Prefetchers (MoP), and continuously refine it through simulation. This approach represents the first practical, agent-driven framework capable of generating synthesizable RTL for prefetchers, outperforming state-of-the-art handcrafted designs on previously unseen workloads. Evaluated on SPEC CPU2006/2017, MoP achieves a 61.1% geometric mean IPC improvement over no prefetching and surpasses Alecto, Berti, and Pythia by 14.5%–23.6%. Its RTL implementation in a 6nm process occupies only 110KB of storage and 0.0347mm² of silicon area.