Score
Design and implement systems that generate synthetic sensor outputs across one or more modalities, including photorealistic renderings and time-series streams. This work includes modeling sensor noise and artifacts, emulating multi-sensor synchronization, and producing real-time GPU-accelerated outputs for testing, evaluation, and analysis.
This work addresses the lack of systematic understanding regarding the success and failure mechanisms of generative models on real-world sensor time-series data. The authors propose SensorGen, the first unified framework for generating and evaluating multi-domain, multimodal sensor signals. They conduct a comprehensive benchmark of five prominent generative model families—including flow matching, diffusion, and autoregressive models—across four domains, seven datasets, and twelve signal modalities, introducing novel techniques for time–frequency modeling and covariate integration. Their findings reveal that flow matching models consistently achieve the best overall performance, that signal characteristics substantially influence generation quality, and that high-fidelity synthetic data can significantly enhance downstream task performance, thereby demonstrating its practical utility.
Existing autonomous driving simulators face two key limitations: insufficient scenario diversity in graphics-based engines (e.g., CARLA) and poor generalizability in learning-based methods (e.g., NeuSim), which are restricted to specific object categories and require dense multi-sensor annotations. To address these bottlenecks, we propose a real2sim2real end-to-end scalable simulation framework. Our method integrates 3D generative modeling, real-to-sim domain translation, forward multi-sensor simulation, and inverse rendering to establish a closed loop: automatically mining rare driving scenarios from real-world data, generating high-fidelity, category-agnostic 3D object assets, and synthesizing corresponding multi-modal sensor data. Crucially, it operates without category priors or dense annotations, significantly improving rare-scenario coverage and data efficiency. Experiments demonstrate that the synthesized data substantially outperforms both conventional computer graphics–based and learning-based baselines in training perception models for robustness.
To address the high cost, environmental constraints, and safety challenges associated with real-world LiDAR data collection, this paper proposes an automated, multimodal synthetic data generation framework built on CoppeliaSim. The framework integrates time-synchronized ToF LiDAR, RGB/depth cameras, and 2D laser scanners to generate high-fidelity point clouds (PCD/PLY) and images (RGB/depth) within urban scenes, accompanied by precise ground-truth pose annotations and timestamps. Notably, it is the first to embed LiDAR-specific security testing—such as adversarial point injection and spoofing attacks—directly into the simulation pipeline, enabling scalable, reproducible, cross-modal dataset construction with fine-grained annotations. Experimental evaluation demonstrates its effectiveness in autonomous driving perception, robotic localization, and LiDAR security vulnerability modeling. The complete codebase, documentation, and animated sample sequences are publicly released.
This work addresses the challenge of interpreting multimodal sensor data for non-experts—stemming from its inherent complexity, cross-modal semantic gap, and dynamic time-varying nature—by proposing the first generative augmented reality (AR) system tailored for domain-agnostic users. Methodologically: (1) it introduces centroid interpolation-based cross-modal embedding, mapping raw sensor data into a pretrained visual latent space; (2) it designs an end-to-end AR scene generation framework requiring no domain-specific knowledge; and (3) it incorporates latent variable reuse and caching to enhance efficiency. Technically, the system integrates cross-modal embedding, foundation model–driven generation, and 3D Gaussian Splatting (3DGS), achieving an 11× reduction in inference latency without compromising visual fidelity. A user study with 485 participants demonstrates statistically significant improvements over baselines in explanation accuracy, cross-scenario consistency, and real-world applicability.
To address the high computational cost and poor generalizability of physics-based models in autonomous driving simulation, this paper presents a systematic survey of data-driven camera and LiDAR simulation methods. It introduces, for the first time, a unified taxonomy of sensor simulation paradigms from two complementary perspectives: generative modeling and neural volume rendering. A novel classification framework for volume renderers is proposed based on input encoding types. The survey identifies two critical challenges: the absence of standardized evaluation protocols and limited cross-scenario generalization. Drawing on over 120 scholarly works, it constructs a comprehensive taxonomy, categorizing generative architectures into five types and volume rendering input encodings into four classes; it further synthesizes six mainstream evaluation metrics alongside their applicability boundaries. The work establishes a theoretical foundation and practical guidance for developing efficient, scalable, multimodal sensor simulation models.
Existing event camera simulators rely on frame-based sequences to infer event timestamps, struggling with fast motion and occlusions, which degrades simulation accuracy. This work proposes a continuous-time event simulator based on dynamic 3D Gaussian splatting that explicitly models per-pixel brightness change rates through a 3D scene representation, enabling precise prediction of threshold-crossing times. The method is the first to generate multiple events within a single rendering step without temporal upsampling. It further incorporates an occlusion-aware adaptive time-stepping scheme and a tile-based arbiter to emulate real sensor bandwidth constraints. Evaluated on RGB–event paired benchmarks, the approach achieves state-of-the-art fidelity in simulated event streams and demonstrates superior transfer performance in downstream tasks.
This work addresses the challenge of cross-modal RGB-X sensor data alignment, which typically relies on expensive hardware calibration. The authors propose a novel cross-modal view synthesis method that operates without requiring depth or calibration information from the X modality. By leveraging only low-cost COLMAP processing on RGB images, the approach achieves 3D-consistent novel view synthesis through a pipeline comprising RGB-X image matching, confidence-aware guided point cloud densification, self-matching filtering, and integration with 3D Gaussian Splatting. This study presents the first demonstration of high-quality cross-modal alignment under the absence of any 3D priors from the X modality, substantially lowering the barrier for multimodal data acquisition and overcoming a key bottleneck in scaling real-world RGB-X dataset collection.
This work addresses the limited adoption of event cameras in robotics due to hardware scarcity and the absence of simulation tools compatible with modern platforms. We present the first multimodal event camera plugin for NVIDIA Isaac Sim, supporting both grayscale and Bayer RGGB modes while synchronously outputting RGB, APS frames, event streams, depth, and IMU data. To enhance effective temporal resolution without compromising compatibility with Isaac Sim’s rendering pipeline, we introduce a motion-guided inter-frame synthesis strategy. Built on native ROS 2 interfaces and supporting multiple resolutions, the plugin incurs less than 400 MB of additional GPU memory on an RTX 4060, enabling real-time simulation at equivalent event rates of 240–960 Hz, with per-frame generation latency as low as 6.98 ms (grayscale) and 7.58 ms (Bayer RGGB).
Existing underwater 3D sonar simulations predominantly rely on LiDAR-like geometric rendering, neglecting critical acoustic effects such as refraction, multipath interference, and phase dependence, thereby limiting fidelity. This work proposes a modular 3D sonar simulation framework that, for the first time, integrates GPU-accelerated graphics rendering with a physics-based acoustic propagation model within the general-purpose NVIDIA Isaac Sim platform. The system enables voxelized sonar simulation of the Water Linked 3D-15 sensor and incorporates FastLIO2 SLAM alongside multisensor fusion (sonar/DVL/IMU/pressure). Supporting hardware-in-the-loop validation, it provides a scalable foundation for fully acoustics-driven volumetric perception. Experimental results demonstrate qualitative consistency with real-world sheet pile data collected in a harbor environment, while also highlighting persistent gaps between current simulation capabilities and physical reality.