🤖 AI Summary
To address the bottleneck where data generation rates in high-throughput imaging systems (e.g., PRISM) vastly exceed real-time processing capabilities, this work proposes a scalable FPGA-based streaming preprocessing architecture. The method jointly optimizes DRAM access patterns, inter-frame subtraction, and mean filtering, leveraging AXI4 burst transfers and streaming buffers to achieve real-time denoising and compression within a single frame interval. Implemented via high-level synthesis (HLS), the hardware pipeline sustains PRISM-scale throughput (>10 Gbps). Experimental results demonstrate sub-frame end-to-end latency, 3.2× raw data compression, and substantial offloading of subsequent CPU/GPU analytics. The core contributions are a frame-rate-constrained, low-latency on-chip denoising architecture and an efficient DRAM scheduling mechanism optimized for streaming imaging workloads.
📝 Abstract
High-throughput imaging workflows, such as Parallel Rapid Imaging with Spectroscopic Mapping (PRISM), generate data at rates that exceed conventional real-time processing capabilities. We present a scalable FPGA-based preprocessing pipeline for real-time denoising, implemented via High-Level Synthesis (HLS) and optimized for DRAM-backed buffering. Our architecture performs frame subtraction and averaging directly on streamed image data, minimizing latency through burst-mode AXI4 interfaces. The resulting kernel operates below the inter-frame interval, enabling inline denoising and reducing dataset size for downstream CPU/GPU analysis. Validated under PRISM-scale acquisition, this modular FPGA framework offers a practical solution for latency-sensitive imaging workflows in spectroscopy and microscopy.