domain randomization

Generating diverse synthetic variations in simulation (appearance, physics, and scene parameters) to produce training data that improves transfer to the real world. Used to design renderings and data-generation pipelines that yield robust real-world performance across sensors, detectors, and observers.

domainrandomization

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Synthetic Video Enhances Physical Fidelity in Video Synthesis

Mar 26, 2025
QZ
Qi Zhao
🏛️ ByteDance | Peking University | ShanghaiTech University | National University of Singapore

This study addresses the low physical fidelity of generative video models—manifested as artifacts such as jittering and interpenetration—by proposing a physics-aware enhancement method grounded in synthetic video. Methodologically, it employs a differentiable rendering pipeline to generate physically consistent synthetic videos, establishes a physics-perceptive data filtering mechanism, and introduces cross-domain feature alignment coupled with adversarial physical consistency regularization—enabling physics realism transfer without differentiable simulation or explicit physical modeling. This work provides the first empirical evidence that synthetic video can substantially improve physical fidelity in video generation. Evaluated on three physics-sensitive tasks—rigid-body collisions, fluid motion, and pendulum dynamics—the approach reduces physical violation rates significantly, achieving an average 37.2% improvement in physical plausibility, validated jointly by user studies and automated physical violation detection.

Enhancing physical fidelity in video generation modelsLeveraging synthetic videos for 3D consistencyReducing artifacts by transferring physical realism

The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.

AI trainingdata qualitydata volume

This work addresses the significant appearance gap between synthetic and real images—commonly referred to as the sim2real appearance gap—that limits the applicability of synthetic data in real-world vision tasks. The authors propose a hybrid augmentation framework that, for the first time, integrates the geometric and material generation capabilities of the diffusion model FLUX.2-4B Klein with the distribution-matching strengths of the image-to-image translation model REGEN. This combination enhances visual realism while preserving semantic consistency. Experimental results demonstrate that the proposed approach substantially narrows the sim2real appearance gap, outperforming individual models in overall quality, with REGEN contributing notably superior photorealism.

appearance gapgame enginephotorealism

This work addresses the critical gap in understanding whether synthetic images are truly interchangeable with real ones in model training and the absence of systematic evaluation frameworks to ensure their safe and effective use. The study systematically quantifies discrepancies between synthetic and real images across three dimensions: high-dimensional feature distributions, low-level statistical properties in color space, and model training dynamics. Building on these insights, the authors propose a pre-evaluation metric for synthetic data of unknown quality and a safety-aware data fusion strategy for training. Experiments demonstrate that carefully calibrated mixing ratios and integration methods of synthetic and real data can substantially enhance model performance and robustness, thereby offering both theoretical grounding and practical guidance for the reliable deployment of synthetic data in machine learning pipelines.

data qualityimage classificationmodel safety

Latest Papers

What's happening recently
View more

Controllable human video generation is hindered by the scarcity of real-world data, particularly for rare identities and complex motion scenarios. This work proposes a unified diffusion-based framework that systematically investigates, for the first time, the synergistic mechanisms between synthetic and real data in human-centric video generation. It reveals their complementary roles and introduces an efficient synthetic sample selection strategy to enhance training. The proposed approach significantly improves motion realism, temporal coherence, and identity fidelity in generated videos, establishing a new paradigm for building data-efficient and generalizable controllable video generation models.

controllable human video generationdata scarcityhuman-centric video synthesis

Computer vision training dataset generation for robotic environments using Gaussian splatting

Dec 15, 2025
PN
Patryk Niżeniec
🏛️ Nicolaus Copernicus University in Toruń

To address domain shift between synthetic and real-world images and high annotation costs in robotic vision tasks, this paper proposes an automated training dataset generation pipeline tailored for robotic environments. Methodologically, it introduces a novel two-pass rendering framework based on 3D Gaussian Splatting, integrating proxy-mesh shadow mapping with splatting-based image synthesis to achieve physically plausible shadow and highlight modeling. Concurrently, it generates pixel-accurate segmentation masks compatible with mainstream detectors such as YOLO. By training on a hybrid dataset comprising a small set of real images and large-scale, high-fidelity synthetic data, the pipeline significantly improves object detection and instance segmentation accuracy. Experiments demonstrate that the approach effectively bridges the domain gap while maintaining high rendering efficiency, offering a scalable, efficient paradigm for building robust robotic vision models.

Automates labeling to eliminate manual annotation bottlenecksBridges domain gap between synthetic and real-world imageryGenerates realistic synthetic datasets for robotic vision

The Impact of Synthetic Data on Object Detection Model Performance: A Comparative Analysis with Real-World Data

Oct 14, 2025
MB
Muammer Bay
🏛️ LycheeAI | Hochschule Hannover | L3S Research Center | Leibniz University Hannover

This study addresses the limited generalization of object detection models in warehouse logistics due to scarce real-world annotated data. To systematically evaluate the efficacy of synthetic data, we propose a balanced fusion training strategy that jointly fine-tunes YOLO-series detectors on real pallet images and high-fidelity, diverse synthetic warehouse scenes generated via NVIDIA Omniverse Replicator. Our approach maintains strict control over scene semantics, lighting, occlusion, and viewpoint variation. Experiments demonstrate that, while substantially reducing annotation costs, the method improves mean Average Precision (mAP) by 3.2–5.7 percentage points over real-data-only baselines. Moreover, the resulting models exhibit enhanced robustness to occlusion, illumination changes, and viewpoint shifts. These results validate the practical utility and scalability of controllable synthetic data for complex industrial vision tasks.

Assessing balanced synthetic-real data integration for robust detectionComparing synthetic versus real data for object detection model trainingEvaluating cost-effective synthetic data in warehouse logistics applications

This study systematically evaluates the suitability and effectiveness of synthetic data across three canonical scenarios: data sharing, model training augmentation, and variance reduction in statistical estimation. By integrating formal modeling, theoretical analysis of generative models, and empirical case studies, the work presents the first comprehensive taxonomy of synthetic data applications and delineates their boundaries of applicability. The research elucidates both the potential and fundamental limitations of synthetic data in enhancing privacy preservation, model performance, and statistical stability. It further demonstrates that many existing or proposed use cases are misaligned with the intrinsic properties of synthetic data, thereby providing decision-makers with a principled theoretical framework to assess whether synthetic data is appropriate for addressing specific data availability challenges.

data augmentationdata sharingprivacy

This work proposes a GAN-inspired privacy-preserving synthetic data generation method that avoids direct access to original data during training. Instead, it leverages fuzz testing to produce candidate samples and iteratively refines them through a discriminator-guided feedback loop combined with statistical distribution constraints to approximate the original data distribution. By innovatively integrating fuzz testing, adversarial discrimination, and indirect constraint mechanisms, the approach achieves strong privacy guarantees—effectively resisting membership inference and data reconstruction attacks—while preserving high data utility. Extensive experiments on four benchmark datasets demonstrate that the proposed method strikes a superior balance between privacy protection and data fidelity compared to existing techniques.

data confidentialityprivacy preservationstatistical distribution

Hot Scholars

YG

Yuhong Guo

Professor, Carleton University
Machine LearningNatural Language ProcessingComputer VisionMedical Data Analysis
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
TL

Tingwen Liu

Institute of Information Engineering, Chinese Academy of Sciences
Content SecurityNatural Language ProcessingKnowledge Graph
JS

Jiawei Sheng

Institute of Information Engineering, Chinese Academy of Sciences
Knowledge GraphRecommendationNatural Language Processing
LL

Lingkun Luo

Shanghai Jiaotong University
Computer visionMachine learningOptimization