Score
Designs, implements, and maintains simulation environments and frameworks that represent dynamic environments over time (including grid-worlds, tournaments, and other task environments), develop and integrate simulators, and support simulation-to-real transfer. Builds scalable, parallel and asynchronous orchestration and data-collection pipelines to run large- or massive-scale simulations, manage distributed execution, and integrate simulators with real systems for training, benchmarking, and data collection.
This study addresses the lack of systematic synthesis at the intersection of artificial intelligence (AI) and modeling and simulation (M&S) by proposing, for the first time, a structured framework based on the full M&S lifecycle—encompassing model construction, input modeling, execution, experimentation, validation, and output analysis. It elucidates the bidirectional integration mechanisms between AI and simulation: how AI enhances or substitutes traditional simulation components, and how simulation supports AI training and evaluation. Incorporating generative AI technologies such as large language models, the paper identifies representative application paradigms and integration approaches across each phase, synthesizes key achievements, and presents a conceptual roadmap tailored to the rapidly evolving ecosystem, while highlighting current limitations and open research challenges.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
Large-scale, heterogeneous metadata in scientific simulations impede result reproducibility and cross-team sharing. To address this, we propose a hardware- and software-agnostic, user-definable two-stage metadata governance framework: (1) non-intrusive acquisition of raw metadata, and (2) on-demand, dynamic structuring. Our key contribution is the first lightweight, general-purpose metadata governance paradigm that decouples acquisition from structuring, enabling zero-code integration into existing HPC simulation workflows. Implemented via the Python-based tool Archivist, the framework supports dynamic schema mapping, declarative configuration, and HPC-adapted interfaces. Evaluated in neuroscience and hydrology simulation use cases, it significantly improves metadata completeness, queryability, and cross-team sharing efficiency—thereby strengthening reproducible and sustainable numerical experimentation.
This work addresses the current lack of open-source infrastructure capable of efficiently training and evaluating large-scale agents on complex tasks such as software engineering and computer operation. To this end, we propose a three-service decoupled architecture tailored for agent-environment interaction workloads, which separates the system into three independent services—model, agent, and environment—enabling fine-grained task scheduling, dynamic resource allocation, and unified interface communication. This design allows each component to scale independently and configure resources flexibly, significantly improving training efficiency and resource utilization. Experimental results demonstrate that the system can stably support tens of thousands of concurrent agent tasks, thereby filling a critical gap in infrastructure for large-scale agent training.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
This work addresses the orchestration bottlenecks faced by ultra-large-scale Sim-AI workflows on leadership-class supercomputers, which arise from task heterogeneity and extreme ensemble sizes. To overcome these challenges, the authors propose EnsembleLauncher, a recursively hierarchical and fully decentralized workflow orchestrator that introduces a decentralized control plane and a programmable scheduling policy interface, thereby surpassing conventional tools in both scalability and scheduling flexibility. Experiments on the Aurora supercomputer demonstrate that EnsembleLauncher can efficiently schedule system-wide resources to support up to 8 million serial tasks, achieving more than a fourfold performance improvement over state-of-the-art alternatives. Furthermore, it significantly enhances resource utilization for workloads with high task variance and active learning pipelines.
Existing cluster-level full-stack simulation struggles to simultaneously achieve high fidelity and high performance. This work proposes the concept of a “simulation-native operating system,” which integrates simulation control and orchestration into the OS kernel, thereby constructing a full-stack simulation framework built upon the Linux virtualization stack. The framework employs four key mechanisms—simulation-oriented scheduling, real-time memory hierarchy management, simulation-aware inter-process communication (IPC), and distributed simulation orchestration—to seamlessly co-execute real and simulated components without requiring modifications to production systems. Experimental results demonstrate that this approach significantly enhances the performance and configuration exploration efficiency of large-scale cluster simulations while preserving full-stack fidelity.
This work addresses the performance bottleneck in large-scale materials screening caused by single-agent architectures and sequential tool invocation, which fail to leverage the parallel capabilities of supercomputers. To overcome this limitation, the authors propose a hierarchical multi-agent framework tailored for supercomputing environments: a central planning agent dynamically partitions tasks, while parallel execution agents collaboratively process them. Efficient orchestration is achieved through integration with a shared Model Context Protocol server and the Parsl workflow engine. This approach establishes the first scalable multi-agent orchestration paradigm that breaks the parallelization barrier of large language models in scientific automation. Demonstrated on the Aurora supercomputer using the gpt-oss-120b model, the system successfully performs high-throughput screening of atmospheric water harvesting candidates from the CoRE MOF database, achieving high task completion rates with minimal orchestration overhead.
Existing evaluation methods struggle to disentangle the quality of task orchestration in multi-agent systems from confounding factors such as agent capabilities and environmental noise, while real-world execution incurs prohibitive costs. To address this, this work proposes OrchBench—a deterministic simulation-based benchmarking platform that models task dependencies via directed acyclic graphs and enables isolated, efficient assessment of orchestration plans. OrchBench achieves the first interpretable evaluation of orchestration quality with dramatically reduced overhead: requiring only 1.3% of the tokens and 10.3% of the time compared to real execution, while maintaining high fidelity (Pearson r = 0.816). Furthermore, it reveals that information retention rate is more critical to performance than simply increasing the number of agents.
This work addresses the challenge of non-reproducible simulation results caused by scheduling dependencies inherent in browser asynchronous runtimes—such as the event loop, Web Workers, and rendering callbacks. To resolve this, the authors propose a logical state visibility stabilization mechanism that decouples asynchronous execution from externally observable state, exposing outputs only after logical state stabilization. This approach ensures consistency between serial and parallel executions without requiring execution replay or explicit scheduler control. Building upon this mechanism, the authors implement a deterministic parallel simulation framework within a real browser environment, effectively restricting asynchronous tasks and rendering callbacks from accessing intermediate states. The framework achieves bit-identical output across varying Worker configurations, scheduling policies, and rendering frequencies, demonstrating both effectiveness and scalability.