HeteroReason: Heterogeneous FPGA-GPU Acceleration for Disaggregated Speculative Reasoning

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient robustness of forward trajectories and low utilization of homogeneous GPU resources in speculative decoding for large reasoning models by proposing an FPGA-GPU heterogeneous algorithm-hardware co-design. Methodologically, the draft model is offloaded to an FPGA while verification is deployed on a GPU. A step-level backtracking mechanism is introduced to enhance robustness, and a parallel pipeline is achieved through shadow synchronization and lookahead speculation scheduling, effectively decoupling the prefill and decoding phases. Experimental results demonstrate that the proposed approach improves average accuracy by 4.2%, achieves latency speedups of 1.01× to 1.42×, and increases energy efficiency by 1.25× to 1.57×.
📝 Abstract
Large Reasoning Models (LRMs) have achieved state-of-the-art performance in reasoning tasks by utilizing Chain-of-Thought (CoT) reasoning. To achieve fast execution speed, speculative reasoning techniques adopt a lightweight draft model for candidate token generation followed by process reward models (PRMs) for verification and a strong target model for refinements. This paper identifies that the existing speculative reasoning paradigm follows a strictly forward-only reasoning trajectory, which lacks robustness and can lead to severe error propagation if early reasoning steps are suboptimal. Furthermore, executing these disparate inference schemes, including sequential drafting and parallel verification on homogeneous GPU platforms, can lead to severe resource underutilization. To address this, we propose HeteroReason, an algorithm-hardware co-designed heterogeneous FPGA-GPU inference paradigm specifically tailored for LRM speculative reasoning. At the algorithmic level, we introduce a backtracking-enhanced workflow that enables the system to recover from low-quality states and explore alternative reasoning trajectories, significantly improving reasoning robustness. At the system level, the draft model is offloaded to the FPGA while deploying the PRM and target models on GPUs. A specialized workflow is optimized to achieve prefill-decode disaggregation, which exploits shadow synchronization to overlap GPU-side refinements with FPGA-side token updates to effectively hide synchronization latency. To mitigate inherent sequential constraints, we propose a step-ahead speculation and refinement scheduling scheme, transitioning the system from a sequential execution scheme to a parallel pipeline. Experimental evaluations show an average 4.2% accuracy improvement, with 1.01x-1.42x latency speedups and 1.25x-1.57x improvements in energy efficiency compared to homogeneous GPU baselines.
Problem

Research questions and friction points this paper is trying to address.

Speculative Reasoning
Large Reasoning Models
Error Propagation
Resource Underutilization
Heterogeneous Acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Reasoning
FPGA-GPU Heterogeneous Acceleration
Backtracking-enhanced Workflow
Prefill-Decode Disaggregation
Algorithm-Hardware Co-design
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zehuan Zhang
Imperial College London
Q
Quan Deng
Imperial College London, Tsinghua University
Z
Zibo Ren
Imperial College London
H
Hao Mark Chen
Imperial College London
G
Guoyu Li
Imperial College London
X
Xuchun Hu
Imperial College London
J
Jose G. F. Coutinho
Imperial College London
Ce Guo
Ce Guo
Imperial College London
Reconfigurable computingRisk Management
Wayne Luk
Wayne Luk
Professor of Computer Engineering, Imperial College London
Hardware and ArchitectutreReconfigurable ComputingDesign Automation
Z
Zhiqiang Que
Imperial College London, University of Bristol
H
Hongxiang Fan
Imperial College London