Score
Designs, implements, and evaluates algorithms and processing pipelines that ingest digital images and produce processed images, measurements, or derived representations through operations such as filtering, denoising, enhancement, geometric and photometric transformations, segmentation, registration, and feature extraction. Works with image formats, acquisition artifacts, sampling and resolution issues, and performance metrics to analyze image quality, algorithmic robustness, and computational resource requirements.
This work addresses the challenge of achieving real-time performance, high accuracy, and energy efficiency in embedded vision systems operating on resource-constrained hardware. The authors propose an algorithm-hardware co-design methodology tailored for DSP/FPGA platforms, optimizing edge, corner, and blob detection operators through hardware-aware algorithmic refinements and quantization techniques. To further enhance throughput without compromising image quality, the approach incorporates inter-frame redundancy elimination and adaptive frame averaging strategies. Experimental results demonstrate that, compared to conventional solutions, the proposed method delivers significantly improved processing speed and energy efficiency, enabling scalable and highly effective real-time embedded vision across diverse applications such as automotive systems, surveillance, and robotics.
This work addresses the poor real-time performance and low energy efficiency of compute-intensive pixel-level operations—such as median filtering and color space conversion—in embedded edge vision applications. We propose a co-optimization methodology based on software-configurable processor (SCP) arrays. Our approach enables, for the first time, real-time median filtering of DVR-standard video streams on a single SCP; furthermore, we introduce a novel inter-processor collaborative execution scheme for color space conversion across an SCP array, overcoming the inflexibility of ASICs and the programming complexity of FPGAs. Through co-design of hardware bitstream compilation, parallel pixel-processing architecture, and on-chip reconfigurable logic, our implementation achieves real-time frame rates for median filtering with 42% lower power consumption, and improves color conversion throughput by 3.8×. Experimental results demonstrate the SCP array’s significant advantages in delivering high-efficiency, highly configurable edge vision computing.
To address perception distortion induced by sensor data compression and virtualization in autonomous driving, this paper proposes a four-step quantitative framework: (1) constructing paired distorted image sets; (2) measuring image fidelity degradation using LPIPS, SSIM, and PSNR; (3) evaluating task-level performance degradation—specifically mAP reduction and increased localization error—on object detection models (YOLO, Faster R-CNN); and (4) establishing statistical correlations between image quality metrics and downstream task performance. This work is the first to quantitatively characterize the relationship between image distortion magnitude and model robustness degradation in autonomous driving perception tasks. Results show that LPIPS exhibits the strongest correlation with performance degradation (Spearman’s ρ > 0.92), significantly outperforming SSIM and PSNR. The framework establishes a reproducible, interpretable, data-quality-driven paradigm for robustness validation of machine learning systems in safety-critical perception applications.
This study addresses the limitations of conventional image enhancement, filtering, and pattern recognition—namely, heavy reliance on manual feature engineering and insufficient real-time performance—by proposing a theory-driven, end-to-end machine learning framework. Methodologically, it is the first to systematically integrate discrete Fourier transform (DFT), Z-transform, and continuous Fourier analysis into deep learning pipelines, synergistically coupling them with convolutional neural networks (CNNs) and classical digital filtering algorithms to enable frequency-domain-guided automated feature extraction and real-time joint signal–image processing. The key contributions include: (i) development of an extensible Python framework; (ii) average PSNR improvement of 3.2 dB in image enhancement and noise suppression tasks; and (iii) 40% acceleration in feature extraction efficiency. This work establishes a novel paradigm for AI-powered real-time computer vision that simultaneously ensures high performance and interpretability.
Whole-slide image (WSI) quality control (QC) faces challenges including low segmentation accuracy for blur, fold, pen-mark, and tissue regions, alongside high annotation costs. Method: We propose a lightweight multi-task semantic segmentation pipeline featuring a collaborative, lightweight CNN/U-Net variant; HistoROI-guided automatic patch-level annotation; pre-screening via patch classification; multi-scale feature fusion; and GPU inference optimization. Contribution/Results: To our knowledge, this is the first method enabling pixel-wise joint segmentation of all four QC regions. Evaluated on over 11,000 WSIs across all 28 TCGA organ sites, it demonstrates strong generalizability and significantly outperforms conventional approaches across all metrics. We publicly release the trained models, source code, annotated datasets, and evaluation results—enabling plug-and-play deployment and domain adaptation.
研究比较了五个库的七种输入路径,从RGB JPEG文件到CUDA float16批处理,评估图像增强管道的吞吐量和GPU内存使用,以优化数据预处理效率。
This work addresses the cumbersome and inefficient model selection and hyperparameter tuning processes that hinder practical deployment in 3D biomedical image analysis. We propose a two-stage Bayesian optimization–driven automated pipeline that jointly optimizes segmentation models, post-processing parameters, and classifier architectures along with pretraining strategies. A key innovation is the introduction of a segmentation-quality–oriented evaluation metric as the optimization objective, coupled with an auxiliary pseudo-labeling mechanism derived from segmentation outputs to substantially reduce manual annotation effort. By integrating domain-adapted synthetic benchmark data, encoder–classifier head architecture search, and prior knowledge–guided pretraining, our method efficiently identifies optimal configurations across four case studies, demonstrating both effectiveness and practical utility.
This study investigates how adolescents informally audit generative AI filters in everyday use to bridge the gap between lived experience and formal algorithmic literacy. Drawing on ethnographic observation and analysis of TikTok videos, and integrating methods from human-computer interaction and learning sciences, the research reveals that high school students spontaneously employ strategies—such as rapid testing, adjusting facial expressions, and manipulating camera angles—to systematically probe the limitations of AI filters. This work is the first to demonstrate the sophisticated algorithmic auditing capabilities adolescents exhibit in unsupervised contexts. It proposes a “hybrid design” approach that integrates everyday practices with formal education, offering an empirical foundation for designing AI literacy curricula and algorithmic transparency mechanisms tailored to youth.
Image degradation is pervasive throughout the imaging pipeline, yet existing research lacks a unified taxonomy and evaluation protocol, hindering cross-dataset and cross-task comparisons. This work introduces a causal perspective to address this gap, proposing a dual-axis classification framework: one axis categorizes degradations by their dominant causal source in the imaging pipeline—encompassing environment, sensor/optics, ISP/codec, and transmission systems—while the other characterizes their perceptual effects, augmented with a lightweight severity quantification layer. Built upon this framework, the COCO Degradation benchmark leverages PSNR, SSIM, and LPIPS to uniformly measure degradation intensity across physical artifacts, algorithmic perturbations, and perceptual distortions, substantially enhancing the evaluation of object detection model robustness under diverse imaging conditions.
This work addresses instruction-based image editing (IIE)—enabling precise and controllable image manipulation through natural language commands. To advance this emerging field, we establish a unified conceptual framework encompassing task formulation, data curation, model architectures, evaluation protocols, and real-world applicability. We further introduce CDD-IIE Bench, the first comprehensive benchmark for IIE, which facilitates multi-dimensional and fine-grained performance diagnosis. By systematically integrating techniques from GANs, diffusion models, autoregressive models, and large language/vision-language models, we conduct an extensive empirical comparison of prominent open-source methods, elucidating their respective strengths and limitations. Our analysis clarifies key evolutionary trajectories in IIE methodologies and provides the community with a standardized evaluation toolkit and actionable directions for future research.