extend simulators

Designs and implements extensions to simulation platforms, including new annotation layers (e.g., semantic class labels, class-agnostic terrain labeling), dataset-generation modules that produce annotated synthetic data, and adapters that ingest or synchronize external real datasets. Validates and optimizes integration for API/data-format compatibility, runtime performance, and annotation fidelity within the simulator pipeline.

extendsimulators

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of real-world data scarcity, high acquisition costs, and privacy sensitivity in multimodal AI training by introducing Simula, a novel framework that pioneers inference-driven synthetic data generation without requiring any seed data. By integrating an agent-based architecture with a controllable generation pipeline, Simula enables fine-grained control over data characteristics and computational resource allocation, substantially enhancing the interpretability and scalability of synthetic data. Through a comprehensive multidimensional evaluation protocol, the framework simultaneously validates both the intrinsic quality of the generated data and its effectiveness in downstream tasks across multiple benchmarks, offering a practical pathway and design paradigm for AI development under data-constrained conditions.

data generationdata scarcitymulti-modal models

A modular and extensible library for parameterized terrain generation

Jun 24, 2025
EW
Erik Wallin
🏛️ Umeå University

Existing terrain generation tools prioritize artistic expression and visual realism but lack parametric control, reproducibility, and scriptability—hindering their use in intelligent robotic simulation-driven development, where controllable and explicitly defined terrains are essential. To address this, we propose TerrainGen: a highly modular Python library for procedural terrain generation that integrates rule-based modeling with multi-scale noise synthesis. It enables fine-grained parameterization of physical attributes—including slope, surface roughness, and rock density—and adopts a loosely coupled architecture compatible with Blender for automated rendering and object placement. A declarative configuration interface further simplifies terrain specification. Experimental evaluation demonstrates TerrainGen’s effectiveness in synthetic data generation and perception ground-truth annotation, significantly improving controllability, reproducibility, and deployment efficiency in environment construction. By providing a scalable, programmable infrastructure, TerrainGen advances simulation-based robotics development and machine learning training pipelines.

Generates parameterized terrains for simulation-driven machine developmentOvercomes limited parameterization in existing artist-focused terrain toolsProvides modular Python library for reproducible, scriptable terrain generation

The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.

AI trainingdata qualitydata volume

Existing autonomous driving datasets lack sufficient diversity, coordination, and cross-domain support, limiting their utility for training multi-agent, multi-sensor systems. To address this gap, this work proposes a modular data generation pipeline built upon the AVstack framework and the CARLA simulator, capable of efficiently producing terabyte-scale, ground-truth-annotated multimodal data. The pipeline encompasses perspectives from ground vehicles, aerial platforms, and infrastructure sensors, and supports flexible single- or multi-agent configurations under controllable, complex scenarios. This approach represents the first scalable, cross-domain collaborative data generation methodology for autonomous driving, substantially enhancing the customization, training efficacy, and practical applicability of perception and sensor fusion models in cooperative autonomous systems.

autonomous systemslarge-scale learningmulti-agent

Latest Papers

What's happening recently
View more

This study addresses the lack of domain priors in general-purpose data and the high cost of manual annotation by exploring synthetic data curation strategies for task-specific visual perception. Shifting focus from generating more data to determining what data to generate, this work proposes a task-oriented data curation paradigm. Methodologically, it systematically integrates three complementary paradigms—procedural rendering, physics-based simulation, and generative AI—leveraging limited real seed samples to learn sensor appearance characteristics and construct customized synthetic pipelines. Experiments demonstrate that this hybrid strategy achieves reliable Sim-to-Real transfer in surface defect detection, old photo restoration, and 6DoF pose estimation. These results validate that combining controllable supervision with appearance learning constitutes an effective pathway for enhancing model generalization capabilities.

Data CurationSim-to-Real TransferSynthetic Data

Existing synthetic data generation tools often suffer from complex workflows, inconsistent standards, and limited cross-modal extensibility, hindering their ability to meet the data demands of large language models in specialized domains and low-resource languages. This work proposes a configuration-driven, end-to-end open-source framework that standardizes multi-source data synthesis through a unified and controllable paradigm. Featuring a highly modular architecture, the framework flexibly adapts to diverse tasks and supports high-quality data generation across multiple pathways, modalities, and languages. Integrated with both a graphical user interface and command-line utilities, it significantly lowers the barrier to entry for users. Empirical evaluations demonstrate that the framework effectively balances generation efficiency and data quality across various scenarios, thereby accelerating the practical deployment of synthetic data in model training pipelines.

data scarcitymultilingualmultimodal

This study addresses the challenge of natural language–driven simulation model discovery by systematically investigating the impact of data representation, Transformer-based embedding models, and reranking strategies on retrieval performance. By constructing multimodal model metadata and leveraging standard information retrieval metrics, the work presents the first quantitative evaluation of open-source embedding models for this task. Experimental results demonstrate that the proposed approach achieves strong performance in recall@5 and nDCG@5, with reranking substantially enhancing effectiveness on complex queries. These contributions establish the first benchmark framework for AI-enabled model reusability, composability, and interoperability in simulation model retrieval.

AI-driven retrievalmodel discoverymodel reuse

Training API-calling agents requires large-scale, high-quality trajectory data, yet conventional approaches rely on fully executable environments and pre-populated databases, limiting their scalability. This work proposes the first synthetic data generation framework that operates without access to a real execution environment, leveraging only API specifications. By orchestrating large language models to collaboratively generate tasks, simulate stateful API interactions, and filter trajectory quality, the method establishes an end-to-end synthetic pipeline. It achieves, for the first time, fully environment-free synthesis of API interaction trajectories. Evaluated on the AppWorld and OfficeBench benchmarks, models fine-tuned with this synthetic data demonstrate substantial performance gains, confirming the framework’s effectiveness and scalability.

API-calling agentsenvironment-freescalability bottleneck

This work addresses critical limitations in existing large language model (LLM)-based 3D scene generation methods for agricultural applications, including insufficient domain-specific knowledge, lack of validation mechanisms, and inadequate modularity, which collectively constrain controllability and scalability. To overcome these challenges, we propose a modular multi-LLM pipeline that integrates agricultural domain knowledge, few-shot prompting, retrieval-augmented generation (RAG), and Unreal Engine APIs to automatically construct realistic agricultural simulation environments. The architecture enables intermediate validation, structured data handling, and flexible extensibility, substantially enhancing semantic accuracy and visual fidelity. User studies and expert evaluations demonstrate that the system significantly outperforms manual design in both modeling efficiency and output quality, effectively overcoming the bottlenecks of conventional monolithic models in domain adaptation and controllable generation.

3D scene generationagricultural simulationdomain-specific reasoning

Hot Scholars

SY

Shimeng Yu

Georgia Institute of Technology, Dean's Professor
Non-volatile MemoryRRAMFerroelectric MemoriesIn-Memory Computing
XY

Xintao Yan

Assistant Professor, The University of Hong Kong
Intelligent VehiclesSimulationDriver BehaviorAI Safety
ZL

Zongqing Lu

Peking University | BeingBeyond
Reinforcement learning
AE

Ainaz Eftekhar

PhD Student, University of Washington
Computer visionReinforcement LearningEmbodied AIRobotics
JS

Jordi Salvador

Allen Institute for AI
Computer VisionMachine LearningEmbodied AI