Score
Designs and implements pipelines and frameworks that generate systematically randomized variants of simulated or synthetic training environments (for example textures, lighting, geometry, physics parameters, object placement, and sensor noise) and integrates those variants into model training. Builds tools to apply visual and non-visual randomization at scale and analyzes how different randomization strategies affect model robustness, generalization, and sim-to-real transfer.
Existing tool-augmented agents face a scarcity of high-quality training data for online reinforcement learning (RL); synthetic data typically lacks interactivity and compositional structure. Method: We propose RandomWorld, the first pipeline to programmatically generate tool-use trajectories featuring multi-step interactions and cross-tool compositionality—overcoming the static and isolated nature of conventional synthetic data. Our approach integrates programmable environment modeling, supervised fine-tuning (SFT), and Proximal Policy Optimization (PPO)-based online RL into an end-to-end training framework. Contributions/Results: On the NESTFUL benchmark, our method achieves new state-of-the-art (SOTA) performance on two core metrics. Crucially, downstream task performance scales consistently with synthetic data volume—providing the first empirical evidence that high-performance tool-using agents can be trained exclusively on synthetic data.
Existing terrain generation tools prioritize artistic expression and visual realism but lack parametric control, reproducibility, and scriptability—hindering their use in intelligent robotic simulation-driven development, where controllable and explicitly defined terrains are essential. To address this, we propose TerrainGen: a highly modular Python library for procedural terrain generation that integrates rule-based modeling with multi-scale noise synthesis. It enables fine-grained parameterization of physical attributes—including slope, surface roughness, and rock density—and adopts a loosely coupled architecture compatible with Blender for automated rendering and object placement. A declarative configuration interface further simplifies terrain specification. Experimental evaluation demonstrates TerrainGen’s effectiveness in synthetic data generation and perception ground-truth annotation, significantly improving controllability, reproducibility, and deployment efficiency in environment construction. By providing a scalable, programmable infrastructure, TerrainGen advances simulation-based robotics development and machine learning training pipelines.
To address domain shift between synthetic and real-world images and high annotation costs in robotic vision tasks, this paper proposes an automated training dataset generation pipeline tailored for robotic environments. Methodologically, it introduces a novel two-pass rendering framework based on 3D Gaussian Splatting, integrating proxy-mesh shadow mapping with splatting-based image synthesis to achieve physically plausible shadow and highlight modeling. Concurrently, it generates pixel-accurate segmentation masks compatible with mainstream detectors such as YOLO. By training on a hybrid dataset comprising a small set of real images and large-scale, high-fidelity synthetic data, the pipeline significantly improves object detection and instance segmentation accuracy. Experiments demonstrate that the approach effectively bridges the domain gap while maintaining high rendering efficiency, offering a scalable, efficient paradigm for building robust robotic vision models.
To address low modeling efficiency, limited diversity, and poor cross-platform compatibility in articulated object simulation for robotics, this paper proposes the first end-to-end procedural asset generation framework tailored for articulated object segmentation, generalizable reinforcement learning, and sim-to-real transfer. The method leverages Blender to implement parametric articulated structure modeling—supporting joint constraints and physical plausibility—integrates differentiable rendering for automated semantic annotation, and establishes a unified export pipeline compatible with PyBullet, Isaac Gym, and MuJoCo. We develop dedicated generators for five common articulated object categories (e.g., cabinet doors, drawers, folding chairs). Experiments demonstrate that the generated assets significantly improve semantic segmentation accuracy (+12.3% mIoU), policy generalization across tasks (+18.7% success rate), and sim-to-real transfer performance (+24.1% success rate).
This work addresses the lack of reproducible and quantifiable evaluation benchmarks in existing digital twin generation methods, which often rely on subjective qualitative comparisons. To this end, the paper proposes a synthetic image generation framework based on high-fidelity 3D models and programmable camera poses, enabling systematic quantitative assessment of reconstruction results under known ground-truth parameters. The approach introduces, for the first time, a programmable virtual environment coupled with a ground-truth parameter reference mechanism, integrating procedural trajectory generation, photorealistic rendering, and feature-point triangulation-based reconstruction. This framework establishes the first benchmark for digital twin evaluation that supports reproducible and objective comparisons, significantly enhancing the consistency and scientific rigor of assessments across different generation strategies.
Existing visual simulation platforms suffer from high technical barriers and lack user-friendly, controllable, and interactive environments for non-graphics experts, hindering efficient synthetic data generation, out-of-distribution (OOD) evaluation, and closed-loop agent testing. To address this, we propose LychSim—the first Unreal Engine 5-based simulation framework natively integrating the Model Context Protocol (MCP). LychSim enables language-driven dynamic scene editing and precise pose control through a lightweight Python API, procedural high-fidelity scene generation, and semantically aligned 3D annotations. This framework substantially lowers the entry barrier and has been successfully applied to synthetic data engines, adversarial evaluation in reinforcement learning, and language-guided layout generation. The code and annotated datasets will be open-sourced to foster community advancement.
This work addresses the significant appearance gap between synthetic and real images—commonly referred to as the sim2real appearance gap—that limits the applicability of synthetic data in real-world vision tasks. The authors propose a hybrid augmentation framework that, for the first time, integrates the geometric and material generation capabilities of the diffusion model FLUX.2-4B Klein with the distribution-matching strengths of the image-to-image translation model REGEN. This combination enhances visual realism while preserving semantic consistency. Experimental results demonstrate that the proposed approach substantially narrows the sim2real appearance gap, outperforming individual models in overall quality, with REGEN contributing notably superior photorealism.
This study addresses the optimal allocation of limited real-world measurement time between system identification and domain randomization to enhance sim-to-real transfer performance in robot learning. Through controlled simulation-to-simulation experiments on a pendulum system, the work presents the first quantitative analysis of the trade-off between parameter identification accuracy and the breadth of domain randomization under a fixed real-data budget. The results demonstrate that, under identifiable dynamics, even a small amount of real data used for precise system identification significantly reduces the reality gap. In contrast, broad domain randomization—even when encompassing the true system parameters—fails to match the effectiveness of accurate parameter estimation. These findings reveal a key strategy for efficiently leveraging scarce real-world data and establish a new paradigm for improving sim-to-real transfer.
This study addresses the labor-intensive nature of designing assets, layouts, and parameters for physics simulation scenes, which hinders the rapid generation of executable dynamic scenarios from text. To overcome this, we propose a hierarchical agent pipeline built upon the Genesis engine, featuring a Planner-Writer-Critic architecture that constitutes the first agent framework specifically tailored for simulation. Furthermore, we introduce compact Debug Cards distilled from graphical demonstrations to enable character-specific physical guidance and automated execution repair. Across 42 evaluation tasks, our method consistently surpasses existing baselines in both physical fidelity and visual quality. User preference studies under blind testing conditions demonstrate significant improvements, while the framework effectively supports multimodal downstream applications such as dataset construction.
This study addresses the current lack of systematic integration of image-generating generative AI in modeling and simulation. It presents the first comprehensive exploration of text-to-image generation techniques within this domain, proposing tool-agnostic, transferable principles and establishing a localized, reproducible generation pipeline that combines prompt engineering with simulation output mapping. The proposed approach supports diverse applications—including conceptual model representation, visualization of simulation results, generation of instructional materials, and construction of multi-scale model interfaces—thereby offering practitioners a structured knowledge framework to evaluate and adapt this emerging technology. By doing so, it significantly enhances the visual expressiveness and interactive capabilities of simulation systems.