🤖 AI Summary
This study addresses the high computational cost and energy consumption of multiply-accumulate operations during U-Net segmentation inference by proposing a quantized accelerator supporting Most Significant Digit First (MSDF) arithmetic. Methodologically, an exact negative-value detection and sign-decision mechanism based on output digit streams is designed, which, combined with offline-calibrated approximate pruning, enables dynamic low-bit skipping and early termination. Architecturally, a two-stage grouped processing unit is employed to support INT8 operands and intra-stream bias accumulation. Experimental results demonstrate that, implemented in a 45nm process and evaluated on the BraTS dataset, the proposed accelerator reduces computation cycles by 38.38% while achieving a Dice score of 80.58%, with a per-inference energy consumption of only 0.726 mJ.
📝 Abstract
U-Net inference for brain-tumor segmentation requires billions of multiply-accumulate operations, motivating hardware that can reduce computation dynamically rather than relying only on fixed precision or static model compression. Most-significant-digit-first (MSDF) arithmetic exposes the leading digits of a result during computation, enabling output-dependent decisions before the full value is generated. This paper presents an MSDF accelerator for quantized U-Net segmentation with a two-stage grouped processing element supporting signed INT8 operands and in-stream bias accumulation. Four runtime mechanisms operate directly on the output digit stream: exact early negative detection (END) in ReLU layers, exact sign-only decision making in the segmentation head, calibrated low-order-digit skipping, and calibrated pruning. The two approximate mechanisms are selected offline under an accuracy constraint, while execution requires only lightweight control and does not modify the stored weights. On a residual U-Net trained with nnU-Net for BraTS, the proposed mechanisms reduce digit cycles by 38.38\% while achieving a mean Dice score of 80.58\% on 73 held-out cases, compared with 81.20\% for the floating-point model; the exact mechanisms alone reduce cycles by 18.79\% without altering the quantized output. Synthesized in 45~nm, the processing element operates at 500~MHz, occupies 0.858~mm$^2$, and consumes 0.726~mJ per $192\times192$ patch under switching-activity-annotated power analysis. A projected eight-output accelerator with shared activation delivery achieves 16.6~ms latency and 1.67~mJ per patch.