Score
Designs, implements, and integrates algorithms that operate directly on raw image or video pixel data to produce low-level representations such as filtered/denoised images, edges and gradients, optical flow or disparity maps, and local feature descriptors. Analyzes and optimizes these methods for noise behavior, numerical stability, computational cost, and robustness so they reliably serve as inputs to higher-level vision systems.
Standard convolutions, due to their fixed structure, linearity, and reliance on local averaging, struggle to capture complex image characteristics such as low-rank structures, adaptive basis representations, and non-uniform spatial dependencies. This work proposes a unified taxonomy encompassing five classes of structured operators—decomposition-based, adaptive weighting, basis-adaptive, integral/kernel-based, and attention-based—and systematically analyzes their differences along key dimensions including locality, linearity, and equivariance. By leveraging techniques such as singular value/tensor decomposition, content-adaptive weighting, learnable analysis bases, position-dependent nonlinear kernels, and attention mechanisms, the study comprehensively evaluates the performance of these operators across image-to-image and image-to-label tasks. The findings clarify the respective strengths and limitations of each operator class, offering both theoretical insights and practical guidance for future research.
This work addresses the challenge of achieving real-time performance, high accuracy, and energy efficiency in embedded vision systems operating on resource-constrained hardware. The authors propose an algorithm-hardware co-design methodology tailored for DSP/FPGA platforms, optimizing edge, corner, and blob detection operators through hardware-aware algorithmic refinements and quantization techniques. To further enhance throughput without compromising image quality, the approach incorporates inter-frame redundancy elimination and adaptive frame averaging strategies. Experimental results demonstrate that, compared to conventional solutions, the proposed method delivers significantly improved processing speed and energy efficiency, enabling scalable and highly effective real-time embedded vision across diverse applications such as automotive systems, surveillance, and robotics.
Real-world sensor noise exhibits high diversity, making it challenging to jointly optimize raw-image denoising, detail recovery, and compression. Method: We introduce RawNIND—the first paired raw-image dataset—and propose two end-to-end joint denoising-and-compression frameworks operating directly in the raw domain: a Bayer-domain method for computational efficiency and a linear RGB-domain method for cross-sensor generalizability. Our approach unifies denoising, demosaicking, and compression within a single differentiable pipeline, adopting a dual-input-stream deep architecture with rate-distortion-optimized coding. Contribution/Results: This work establishes the “raw-data-first” paradigm for image processing. Experiments demonstrate consistent superiority over sRGB-domain baselines across cross-sensor generalization, PSNR/SSIM, and rate-distortion performance; encoding efficiency improves significantly, and inference speed increases by 40%.
This study addresses the limitations of conventional image enhancement, filtering, and pattern recognition—namely, heavy reliance on manual feature engineering and insufficient real-time performance—by proposing a theory-driven, end-to-end machine learning framework. Methodologically, it is the first to systematically integrate discrete Fourier transform (DFT), Z-transform, and continuous Fourier analysis into deep learning pipelines, synergistically coupling them with convolutional neural networks (CNNs) and classical digital filtering algorithms to enable frequency-domain-guided automated feature extraction and real-time joint signal–image processing. The key contributions include: (i) development of an extensible Python framework; (ii) average PSNR improvement of 3.2 dB in image enhancement and noise suppression tasks; and (iii) 40% acceleration in feature extraction efficiency. This work establishes a novel paradigm for AI-powered real-time computer vision that simultaneously ensures high performance and interpretability.
Traditional hand-crafted histogram features—such as Local Binary Patterns (LBP) and edge histograms—are incompatible with end-to-end deep learning due to their non-differentiability. To address this, we propose a differentiable histogram layer, enabling the first neuralization and learnability of such features. Methodologically, we design Neural Local Binary Patterns (NLBP) and Neural Edge Histogram Descriptor (NEHD) modules, integrated as differentiable statistical layers within CNNs to support gradient backpropagation and joint optimization. Our core contribution lies in unifying hand-engineered feature design with deep learning paradigms, allowing local statistical priors to be data-drivenly learned and enhanced. Extensive experiments on multiple image classification benchmarks and real-world datasets demonstrate consistent and significant performance gains, validating that neuralized histogram features substantially improve representation capability.
This paper addresses the problem of generating dynamic visual illusion images. We propose a zero-shot diffusion sampling framework that enables perceptually controllable, text-guided synthesis via image component decomposition—specifically into frequency-domain, luminance/chrominance, and motion blur subspaces. Our method integrates multi-scale frequency decomposition, component-conditioned sampling, and composite noise estimation. Key contributions include: (1) the first noise-estimation factorization fusion mechanism, enabling joint modeling of heterogeneous component-wise conditions; (2) unified generation of appearance variations governed by spatial distance, illumination, and motion blur dependencies; and (3) component-level inverse editing and controllable re-synthesis of real-world images. Crucially, our approach achieves high fidelity and photorealism without compromising sharpness or perceptual plausibility, thereby extending the paradigm of controllable spatial-perception illusion generation.
Existing generative models struggle to synthesize physically consistent camera raw images, hindering progress in low-level vision tasks. This work proposes RawGen, the first diffusion framework capable of both text-to-raw image generation and sRGB-to-raw inverse mapping. RawGen leverages the generative priors of large-scale sRGB diffusion models and integrates multi-parameter ISP simulation, a conditional denoiser, and a dedicated decoder to jointly produce physically plausible linear raw images in both latent and pixel spaces. To overcome the limitation of fixed ISP assumptions, the authors construct a many-to-one inverse ISP dataset. Experiments demonstrate that RawGen significantly outperforms existing methods in raw reconstruction quality, and its synthetic data effectively enhances performance on downstream vision tasks.
This work proposes UniISP, the first end-to-end image signal processing (ISP) framework that unifies human visual perception with the requirements of downstream machine vision tasks. Traditional ISP pipelines produce RGB images aligned with human aesthetic preferences but often discard information critical for machine vision, whereas raw-data-pass-through approaches fail to meet human visual expectations. To address this trade-off, UniISP introduces a Hybrid Attention Module (HAM) trained under supervised learning to generate images that simultaneously achieve high perceptual quality and preserve task-relevant information. Additionally, it incorporates feature adapters to efficiently transfer useful representations from the ISP stage to downstream task networks. Extensive experiments across multiple datasets and scenarios demonstrate that UniISP consistently enhances both image aesthetics and task performance, confirming its strong generalization capability and effectiveness.
This work proposes SuperCam, a novel camera architecture designed to address the inefficiency of conventional cameras in resource-constrained settings, where they generate excessive redundant data that hinders downstream visual tasks. SuperCam uniquely integrates adaptive superpixel segmentation directly into the hardware pipeline, enabling online, lightweight data compression and preservation of critical visual information at the point of capture. By synergizing this design with edge computing, the system substantially reduces memory footprint and bandwidth requirements. Experimental results demonstrate that SuperCam consistently outperforms traditional approaches across multiple vision tasks—including semantic segmentation, object detection, and monocular depth estimation—thereby validating its feasibility and superiority for efficient perception under stringent resource limitations.
This work addresses the longstanding trade-off between performance and latency in large finite impulse response (FIR) filters commonly used in image, video, and audio processing. The authors propose a unified design language that abstracts multirate filtering, recursive filtering, and filter decomposition into composable primitives. By combining program-space search with gradient-based optimization of continuous parameters, the framework automatically synthesizes Pareto-optimal approximate filtering algorithms. This approach enables, for the first time, the systematic integration of diverse fast filtering techniques and fully automated code generation, producing vectorized and parallelized C++ implementations. Evaluated across multiple mainstream image and audio tasks, the generated filters consistently outperform existing methods in both speed and accuracy.
This work addresses the limitations of existing deep learning approaches for RAW image denoising, which often neglect classical denoising priors, resulting in overly complex models with limited generalization. To overcome this, we propose the first learnable non-local module that explicitly embeds the classical non-local self-similarity prior into a neural network. Our method integrates multi-scale feature extraction, learnable matching and collaborative filtering, noise-level map conditioning, and joint training on both synthetic and real-world noise. The resulting model achieves performance comparable to state-of-the-art CNN- and Transformer-based methods across multiple benchmarks and real datasets, while significantly reducing parameter count and demonstrating strong cross-sensor generalization capability.