Score
Designs and implements simulation models and pipelines that generate synthetic depth measurements (depth maps or point clouds) reproducing real sensor artifacts such as sparsity, missing returns, quantization, range-dependent and high‑frequency noise, and other modality-specific distortions. Builds and calibrates parameterizations and validation procedures to match these models to real sensor data and to reduce the sim‑to‑real perception gap so algorithms trained in simulation can transfer to hardware (including zero‑shot transfer).
To address the high computational cost and poor generalizability of physics-based models in autonomous driving simulation, this paper presents a systematic survey of data-driven camera and LiDAR simulation methods. It introduces, for the first time, a unified taxonomy of sensor simulation paradigms from two complementary perspectives: generative modeling and neural volume rendering. A novel classification framework for volume renderers is proposed based on input encoding types. The survey identifies two critical challenges: the absence of standardized evaluation protocols and limited cross-scenario generalization. Drawing on over 120 scholarly works, it constructs a comprehensive taxonomy, categorizing generative architectures into five types and volume rendering input encodings into four classes; it further synthesizes six mainstream evaluation metrics alongside their applicability boundaries. The work establishes a theoretical foundation and practical guidance for developing efficient, scalable, multimodal sensor simulation models.
To address the high cost, environmental constraints, and safety challenges associated with real-world LiDAR data collection, this paper proposes an automated, multimodal synthetic data generation framework built on CoppeliaSim. The framework integrates time-synchronized ToF LiDAR, RGB/depth cameras, and 2D laser scanners to generate high-fidelity point clouds (PCD/PLY) and images (RGB/depth) within urban scenes, accompanied by precise ground-truth pose annotations and timestamps. Notably, it is the first to embed LiDAR-specific security testing—such as adversarial point injection and spoofing attacks—directly into the simulation pipeline, enabling scalable, reproducible, cross-modal dataset construction with fine-grained annotations. Experimental evaluation demonstrates its effectiveness in autonomous driving perception, robotic localization, and LiDAR security vulnerability modeling. The complete codebase, documentation, and animated sample sequences are publicly released.
Existing autonomous driving simulators face two key limitations: insufficient scenario diversity in graphics-based engines (e.g., CARLA) and poor generalizability in learning-based methods (e.g., NeuSim), which are restricted to specific object categories and require dense multi-sensor annotations. To address these bottlenecks, we propose a real2sim2real end-to-end scalable simulation framework. Our method integrates 3D generative modeling, real-to-sim domain translation, forward multi-sensor simulation, and inverse rendering to establish a closed loop: automatically mining rare driving scenarios from real-world data, generating high-fidelity, category-agnostic 3D object assets, and synthesizing corresponding multi-modal sensor data. Crucially, it operates without category priors or dense annotations, significantly improving rare-scenario coverage and data efficiency. Experiments demonstrate that the synthesized data substantially outperforms both conventional computer graphics–based and learning-based baselines in training perception models for robustness.
This work addresses the challenge of sim-to-real transfer in industrial visual inspection, where multiple domain gaps arise from discrepancies in sensors, lighting conditions, materials, and defect patterns. The authors propose a unified framework centered on the availability of prior knowledge, systematically integrating three scenarios: CAD-available, CAD-unavailable, and boundary cases, thereby harmonizing CAD-driven pose estimation and CAD-free anomaly detection paradigms. Their approach combines CAD-based rendering, RGB-D simulation, synthetic defect generation, pre-trained features, vision-language priors, and test-time geometric consistency verification. Experiments on T-LESS/BOP, MVTec AD, and VisA benchmarks demonstrate that transfer performance hinges more critically on source distribution design, detector capacity, and minimal real-world calibration than on the quantity of CAD renderings; notably, CAD models at test time effectively enable mask generation, pose refinement, and depth consistency validation.
This study addresses the low physical fidelity of generative video models—manifested as artifacts such as jittering and interpenetration—by proposing a physics-aware enhancement method grounded in synthetic video. Methodologically, it employs a differentiable rendering pipeline to generate physically consistent synthetic videos, establishes a physics-perceptive data filtering mechanism, and introduces cross-domain feature alignment coupled with adversarial physical consistency regularization—enabling physics realism transfer without differentiable simulation or explicit physical modeling. This work provides the first empirical evidence that synthetic video can substantially improve physical fidelity in video generation. Evaluated on three physics-sensitive tasks—rigid-body collisions, fluid motion, and pendulum dynamics—the approach reduces physical violation rates significantly, achieving an average 37.2% improvement in physical plausibility, validated jointly by user studies and automated physical violation detection.
Open-world LiDAR perception validation is hindered by the trade-off between poor controllability of real-world scenarios and physical inaccuracies in simulation. To address this, we propose a physics-informed point cloud recombination method using physical human-shaped targets: high-precision multi-pose, multi-material target point clouds are captured in lab settings using an Ouster OS1-128 LiDAR; these are then registered to 3D meshes and jointly rendered with geometric and intensity attributes before being dynamically fused into real-world road-scene point clouds. The resulting synthetic scenes preserve sensor-level physical fidelity—including material-dependent intensity response—while enabling full scene controllability. This work presents the first systematic recombination of physical target point clouds with field-collected data, supporting fine-grained occlusion modeling and joint algorithm-sensor robustness attribution. Experiments show reconstruction error <2.1% versus ground truth, significantly improving reproducibility of edge cases and credibility of failure root-cause analysis.
This work addresses the lack of reproducible and quantifiable evaluation benchmarks in existing digital twin generation methods, which often rely on subjective qualitative comparisons. To this end, the paper proposes a synthetic image generation framework based on high-fidelity 3D models and programmable camera poses, enabling systematic quantitative assessment of reconstruction results under known ground-truth parameters. The approach introduces, for the first time, a programmable virtual environment coupled with a ground-truth parameter reference mechanism, integrating procedural trajectory generation, photorealistic rendering, and feature-point triangulation-based reconstruction. This framework establishes the first benchmark for digital twin evaluation that supports reproducible and objective comparisons, significantly enhancing the consistency and scientific rigor of assessments across different generation strategies.
This work addresses key challenges in deploying 3D deep learning on edge devices, including the unstructured nature of point clouds, high computational costs of conventional preprocessing, and performance degradation caused by domain gaps between synthetic CAD models and real-world LiDAR data. To bridge this domain discrepancy, the authors propose a sensor-aware physically simulated LiDAR data generation method. Furthermore, they introduce a deterministic Critical Point Layer (CPL) that enables efficient point cloud compression without requiring distance-based sorting. Integrated with an ARM Cortex-A76-optimized lightweight classification network, the system compresses input point clouds from 1,024 to 40–60 points and achieves real-time inference at approximately 50 FPS on a Raspberry Pi 5, attaining a classification accuracy of 88.36%.
Existing underwater 3D sonar simulations predominantly rely on LiDAR-like geometric rendering, neglecting critical acoustic effects such as refraction, multipath interference, and phase dependence, thereby limiting fidelity. This work proposes a modular 3D sonar simulation framework that, for the first time, integrates GPU-accelerated graphics rendering with a physics-based acoustic propagation model within the general-purpose NVIDIA Isaac Sim platform. The system enables voxelized sonar simulation of the Water Linked 3D-15 sensor and incorporates FastLIO2 SLAM alongside multisensor fusion (sonar/DVL/IMU/pressure). Supporting hardware-in-the-loop validation, it provides a scalable foundation for fully acoustics-driven volumetric perception. Experimental results demonstrate qualitative consistency with real-world sheet pile data collected in a harbor environment, while also highlighting persistent gaps between current simulation capabilities and physical reality.
To address domain shift between synthetic and real-world images and high annotation costs in robotic vision tasks, this paper proposes an automated training dataset generation pipeline tailored for robotic environments. Methodologically, it introduces a novel two-pass rendering framework based on 3D Gaussian Splatting, integrating proxy-mesh shadow mapping with splatting-based image synthesis to achieve physically plausible shadow and highlight modeling. Concurrently, it generates pixel-accurate segmentation masks compatible with mainstream detectors such as YOLO. By training on a hybrid dataset comprising a small set of real images and large-scale, high-fidelity synthetic data, the pipeline significantly improves object detection and instance segmentation accuracy. Experiments demonstrate that the approach effectively bridges the domain gap while maintaining high rendering efficiency, offering a scalable, efficient paradigm for building robust robotic vision models.
This work addresses the significant appearance gap between synthetic and real images—commonly referred to as the sim2real appearance gap—that limits the applicability of synthetic data in real-world vision tasks. The authors propose a hybrid augmentation framework that, for the first time, integrates the geometric and material generation capabilities of the diffusion model FLUX.2-4B Klein with the distribution-matching strengths of the image-to-image translation model REGEN. This combination enhances visual realism while preserving semantic consistency. Experimental results demonstrate that the proposed approach substantially narrows the sim2real appearance gap, outperforming individual models in overall quality, with REGEN contributing notably superior photorealism.