Scalable FPGA Framework for Real-Time Denoising in High-Throughput Imaging: A DRAM-Optimized Pipeline using High-Level Synthesis

📅 2025-08-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the bottleneck where data generation rates in high-throughput imaging systems (e.g., PRISM) vastly exceed real-time processing capabilities, this work proposes a scalable FPGA-based streaming preprocessing architecture. The method jointly optimizes DRAM access patterns, inter-frame subtraction, and mean filtering, leveraging AXI4 burst transfers and streaming buffers to achieve real-time denoising and compression within a single frame interval. Implemented via high-level synthesis (HLS), the hardware pipeline sustains PRISM-scale throughput (>10 Gbps). Experimental results demonstrate sub-frame end-to-end latency, 3.2× raw data compression, and substantial offloading of subsequent CPU/GPU analytics. The core contributions are a frame-rate-constrained, low-latency on-chip denoising architecture and an efficient DRAM scheduling mechanism optimized for streaming imaging workloads.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: Medical and Biological ImagingCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Data management and stream processing for Web, mobile and wireless applicationsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSecurity and Privacy: Data transparency and provenance
📝 Abstract
High-throughput imaging workflows, such as Parallel Rapid Imaging with Spectroscopic Mapping (PRISM), generate data at rates that exceed conventional real-time processing capabilities. We present a scalable FPGA-based preprocessing pipeline for real-time denoising, implemented via High-Level Synthesis (HLS) and optimized for DRAM-backed buffering. Our architecture performs frame subtraction and averaging directly on streamed image data, minimizing latency through burst-mode AXI4 interfaces. The resulting kernel operates below the inter-frame interval, enabling inline denoising and reducing dataset size for downstream CPU/GPU analysis. Validated under PRISM-scale acquisition, this modular FPGA framework offers a practical solution for latency-sensitive imaging workflows in spectroscopy and microscopy.
Problem

Research questions and friction points this paper is trying to address.

Real-time denoising for high-throughput imaging data
Optimizing FPGA processing to exceed data generation rates
Reducing latency and dataset size for downstream analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

FPGA-based preprocessing pipeline with HLS
DRAM-optimized burst-mode AXI4 interfaces
Real-time frame subtraction and averaging
🔎 Similar Papers
No similar papers found.