Score
Designs and implements simulation-based evaluation pipelines and benchmarks that estimate, compare, and rank the expected real-world performance of decision-making policies. This includes building simulated policy rollouts, deriving correlational validation metrics and uncertainty estimates for sim-to-real transfer, and selecting or prioritizing high-performing policies for real-world testing.
Agent-based modeling (ABM) lacks systematic, standardized benchmarks for policy evaluation, hindering rigorous assessment of its capabilities in real-world policy analysis. Method: We introduce the first ABM capability benchmark specifically designed for policy evaluation, comprising 20 end-to-end policy modeling scenarios, 65 fine-grained subtasks, and 200 automatically generated tasks. Our hierarchical evaluation framework—structured as Scenario → Subtask → Generated Task—integrates multi-agent modeling, behavioral calibration, expert validation, and automated task generation, grounded in social simulation theory and empirical policy analysis frameworks. Contribution/Results: Experiments reveal that state-of-the-art ABM methods achieve only 24.5%, 15.04%, and 14.5% coverage across the three task categories, exposing a substantial gap between current ABM capabilities and practical policy assessment requirements. This benchmark establishes a reproducible, extensible evaluation infrastructure to guide future research and development in policy-oriented ABM.
Current evaluations of vision-language-action (VLA) policies suffer from a lack of reliable correlation between simulation performance and real-world outcomes, limiting their utility in guiding practical deployment. This work presents the first systematic quantification of sim-to-real alignment across multiple simulation platforms, analyzing consistency in policy ranking, performance correlation, and failure modes under perturbations across diverse tasks. Through extensive multi-platform simulations, policy fine-tuning analyses, post-training data scaling studies, and robustness evaluations, the study identifies key characteristics of high-fidelity simulators and proposes design principles to enhance simulation utility. These findings substantially improve the reliability and practical guidance value of simulation in VLA policy development.
Current vision-driven robotic simulation benchmarks have advanced manipulation research but critically lack evaluation capabilities for sim-to-real transfer. Method: We introduce the first benchmark specifically designed for assessing general-purpose manipulation policies’ sim-to-real transferability—featuring a high-fidelity visual simulation environment, a controllable, incrementally complex task suite, a systematic domain perturbation protocol (e.g., lighting, material, motion blur), and novel cross-domain performance alignment metrics. Contribution/Results: Our framework is the first to jointly model task difficulty, perturbation type, and inter-domain performance gap, thereby explicitly exposing generalization bottlenecks of existing policies under realistic conditions. Experiments demonstrate the benchmark’s reproducibility, scalability, and diagnostic utility, establishing a standardized, rigorous testbed for evaluating and improving the robustness and transfer efficiency of general-purpose robotic policies.
This work proposes a novel “betting”-based methodology for efficiently and accurately evaluating robotic performance in real-world settings where physical experimentation is constrained. By introducing betting theory into sim-to-real performance assessment—a first in the field—the approach constructs an estimator theoretically superior to Monte Carlo estimation. The method integrates control variate approximation, cross-fidelity simulation, and statistical decision rules to enable practical deployment. Experimental results demonstrate its efficacy on synthetic data and simulated environments, and it successfully infers real-world robotic grasping accuracy with significantly fewer physical trials while improving evaluation precision.
Traditional benchmarks provide only aggregate scores, offering insufficient evidence to support reliable deployment decisions and thereby creating a disconnect between evaluation and action. To address this gap, this work proposes a “deployment-completeness” benchmarking framework, introducing novel metrics—evidence fibers, completeness curves, and certifiable proportions—alongside a systematic audit methodology comprising evidence fiber analysis, response ranking intervals, conformal coverage evaluation, and a certify-then-acquire decision pipeline. Empirical evaluation on benchmarks such as Tox21, Matbench, and JARVIS reveals that conventional approaches suffer a drastic drop in channel coverage to 10.07% under real-world deployment conditions. In contrast, the proposed method reduces error-driven deployment decisions to 0.027% on Tox21 and 0.128% on JARVIS, substantially enhancing deployment reliability.
This study addresses the sim-to-real gap arising from calibration data confusion and drift in pretrained simulators by investigating how to effectively integrate cheap but biased simulations with expensive yet unbiased real-world experiments in sequential decision-making. The authors extend the simulation lemma to decompose policy value error, revealing that under passive learning, the reachability gap is irreducible. To overcome this limitation, they propose Fisher-SEP, an active experimentation strategy based on the Fisher information matrix, which minimizes the predictive variance of the target policy’s value through Bayesian posterior inference. Empirical validation on two real-world domains—vending machine supply chains and mobile HIV testing—demonstrates that early real-world trials yield substantial long-term benefits in the former, while only active exploration effectively covers low-monitoring regions in the latter.
This work addresses the lack of a unified, reproducible benchmark for evaluating inventory management policies due to heterogeneous environmental assumptions. To this end, we introduce gym-invmgmt, an open-source framework built on Gymnasium that standardizes 22 core scenarios and their multi-agent extensions by harmonizing state transitions, action constraints, reward functions, and key performance indicators. This enables, for the first time, auditable comparisons among optimization-based, heuristic, and learning-based policies under a consistent evaluation protocol. Through systematic experiments, we assess diverse approaches—including stochastic programming, PPO-Transformer, Residual RL, graph neural networks (GNNs), imitation learning, and constrained large language models—revealing that policy performance jointly depends on information access, demand dynamics, network topology, and representation design. Stochastic programming achieves optimal performance at high computational cost, PPO-Transformer offers efficient inference with high policy quality, Residual RL demonstrates robustness, and GNNs excel in divergent topologies but exhibit fragility in serial structures.