Score
Designs and implements simulation pipelines that synthesize depth-dependent degradations in imagery or volumetric sensor data by modeling forward scattering, particle populations, and staged foreground/background effects. Builds parametric and procedural forward models to generate synthetic datasets where degradation strength varies with scene depth and across sequential processing stages for training, evaluation, or analysis of restoration and perception algorithms.
To address the high computational cost and poor generalizability of physics-based models in autonomous driving simulation, this paper presents a systematic survey of data-driven camera and LiDAR simulation methods. It introduces, for the first time, a unified taxonomy of sensor simulation paradigms from two complementary perspectives: generative modeling and neural volume rendering. A novel classification framework for volume renderers is proposed based on input encoding types. The survey identifies two critical challenges: the absence of standardized evaluation protocols and limited cross-scenario generalization. Drawing on over 120 scholarly works, it constructs a comprehensive taxonomy, categorizing generative architectures into five types and volume rendering input encodings into four classes; it further synthesizes six mainstream evaluation metrics alongside their applicability boundaries. The work establishes a theoretical foundation and practical guidance for developing efficient, scalable, multimodal sensor simulation models.
To address the high cost, environmental constraints, and safety challenges associated with real-world LiDAR data collection, this paper proposes an automated, multimodal synthetic data generation framework built on CoppeliaSim. The framework integrates time-synchronized ToF LiDAR, RGB/depth cameras, and 2D laser scanners to generate high-fidelity point clouds (PCD/PLY) and images (RGB/depth) within urban scenes, accompanied by precise ground-truth pose annotations and timestamps. Notably, it is the first to embed LiDAR-specific security testing—such as adversarial point injection and spoofing attacks—directly into the simulation pipeline, enabling scalable, reproducible, cross-modal dataset construction with fine-grained annotations. Experimental evaluation demonstrates its effectiveness in autonomous driving perception, robotic localization, and LiDAR security vulnerability modeling. The complete codebase, documentation, and animated sample sequences are publicly released.
This work addresses the degradation of visual quality in underwater images captured in real oceanic environments, which is primarily caused by depth-dependent forward scattering blur and marine snow artifacts. To tackle this challenge, the authors propose a staged image enhancement framework that explicitly models the depth-dependent forward scattering effect for the first time and extracts realistic marine snow degradation patterns from authentic underwater imagery. These components are leveraged to generate high-fidelity synthetic data for fine-tuning a Joint-ID network, followed by a lightweight contrast enhancement post-processing step. The approach effectively bridges the domain gap between synthetic and real underwater images, yielding significant improvements in UIQM scores and perceptual clarity on a real-world dataset collected off the coast of Korea, thereby enhancing the usability of underwater imagery for robotic operations.
Existing image degradation synthesis models suffer from poor generalizability, reliance on hand-crafted parameters, and support for only a limited set of predefined degradation types—thus failing to capture the complex, realistic mix of homogeneous (global) and heterogeneous (spatially varying) degradations observed in practice. To address this, we propose the first universal, parameter-free image degradation model. Our method achieves unsupervised disentanglement of image content and degradation features for the first time; introduces a compression-driven disentanglement mechanism; and incorporates two novel modules explicitly modeling spatially non-uniform degradations. Built upon an end-to-end autoencoder architecture, the model enables automatic synthesis of diverse, high-fidelity degradations without manual intervention. Evaluated on film grain simulation and blind image restoration, our approach significantly improves degradation realism and downstream task generalization. It delivers plug-and-play, universal degradation synthesis—setting a new standard for data-efficient, realistic image corruption modeling.
This work addresses the degradation of image visibility and multi-view consistency caused by haze, which severely compromises novel view synthesis quality. To tackle this challenge, the authors propose a multi-stage optimization framework that sequentially integrates image restoration, dehazing, enhancement via multimodal large language models (MLLMs), and joint optimization of 3D Gaussian Splatting (3DGS) with Markov Chain Monte Carlo (MCMC), followed by averaging across multiple refinement rounds to simultaneously enhance visibility and preserve cross-view scene consistency. By synergistically combining generative priors with geometric optimization, the method achieved first place among 14 teams in Track 2 of the NTIRE 2026 3DRR Challenge, demonstrating superior quantitative performance and visual quality on the official benchmark.
Remote sensing simulation faces challenges including heavy reliance on LiDAR data, extensive manual intervention, and difficulty in cross-spectral band modeling. Method: This paper proposes an end-to-end, physically grounded simulation framework driven solely by commercial satellite imagery. It automatically constructs 3D geometry from digital surface models (DSMs), infers material properties by fusing multi-source satellite imagery, and integrates physics-based rendering with radiative transfer modeling across a broad spectral range (200 nm–20 μm) to achieve fully automated, high-fidelity reconstruction of terrain, buildings, vegetation, and dynamic vehicles. Contribution/Results: To our knowledge, this is the first framework to eliminate LiDAR dependence and generate novel regional scenes without any manual intervention. Experiments demonstrate substantial reduction in modeling cost, full-spectrum support—from ultraviolet (UV) to long-wave infrared (LWIR)—for algorithm development and processing pipeline validation, and significant improvements in realism, generalizability, and scalability of geospatial scene simulation.
This work addresses the challenges of stereo image restoration in complex degradation environments—such as underwater, hazy, and low-light conditions—where diverse physical degradations and severe information loss hinder performance. Existing datasets are often limited to a single degradation type or lack stereo consistency. To bridge this gap, we introduce M3D-Stereo, a high-resolution dataset comprising 7,904 stereo image pairs captured across multiple media, encompassing four degradation types at six progressive severity levels, all accompanied by pixel-aligned clean ground truth. M3D-Stereo is the first to enable realistic modeling of multi-medium, multi-degradation, and multi-level distortions while preserving stereo consistency, supporting both single-level and mixed-level restoration tasks. Leveraging controlled laboratory acquisition and high-precision alignment, the dataset significantly enhances the fidelity and reliability of evaluating image restoration and stereo matching algorithms in complex scenarios. The dataset is publicly released under the LGPLv3 license.
This work addresses the significant appearance gap between synthetic and real images—commonly referred to as the sim2real appearance gap—that limits the applicability of synthetic data in real-world vision tasks. The authors propose a hybrid augmentation framework that, for the first time, integrates the geometric and material generation capabilities of the diffusion model FLUX.2-4B Klein with the distribution-matching strengths of the image-to-image translation model REGEN. This combination enhances visual realism while preserving semantic consistency. Experimental results demonstrate that the proposed approach substantially narrows the sim2real appearance gap, outperforming individual models in overall quality, with REGEN contributing notably superior photorealism.
Image degradation is pervasive throughout the imaging pipeline, yet existing research lacks a unified taxonomy and evaluation protocol, hindering cross-dataset and cross-task comparisons. This work introduces a causal perspective to address this gap, proposing a dual-axis classification framework: one axis categorizes degradations by their dominant causal source in the imaging pipeline—encompassing environment, sensor/optics, ISP/codec, and transmission systems—while the other characterizes their perceptual effects, augmented with a lightweight severity quantification layer. Built upon this framework, the COCO Degradation benchmark leverages PSNR, SSIM, and LPIPS to uniformly measure degradation intensity across physical artifacts, algorithmic perturbations, and perceptual distortions, substantially enhancing the evaluation of object detection model robustness under diverse imaging conditions.
This study addresses the sensitivity of high-resolution range profiles (HRRPs) to data acquisition conditions, which undermines the robustness of radar automatic target recognition in complex operational scenarios. To tackle this challenge, the authors leverage a large-scale real-world maritime dataset and, for the first time, treat ship geometric parameters—such as physical dimensions and aspect angles—as core conditioning variables to develop a controllable HRRP generative model. This approach overcomes the limitations of prior methods that rely on small-scale or scenario-specific data, successfully reproducing the line-of-sight geometry–driven distribution patterns observed in real HRRPs. The results demonstrate that incorporating geometric conditions is pivotal for enhancing both the fidelity of synthetic HRRPs and their generalization across diverse operational environments.