Score
Defining and applying evaluation metrics and validation protocols that measure robustness and generalization of navigation systems across simulation and real-world robot deployments, ensuring improvements transfer between environments.
Contemporary AI systems are rapidly advancing toward transformative capabilities, necessitating safety evaluation methodologies that transcend conventional static benchmarks. Method: We propose a novel three-dimensional safety assessment framework—“Capability–Propensity–Control”—that systematically integrates measurement targets (e.g., deception capability, power-seeking propensity, adversarial robustness), measurement modalities (behavioral testing, internal analysis), and governance mapping. This framework overcomes limitations of static benchmarking by enabling dynamic, multi-layered evaluation. Contribution/Results: We introduce the first unified taxonomy covering the full stack of AI safety assessment; identify critical evaluation pitfalls—including “safety washing” and “model sandbagging”; formally define safety-critical capabilities and hazardous propensities; provide practitioners with actionable assessment guidelines; establish decision-support interfaces for regulators; and uncover fundamental research gaps in scalable, interpretable, and governance-aligned safety evaluation.
This paper addresses safety risks confronting embodied navigation systems in dynamic real-world environments. We present the first systematic survey spanning three dimensions: attack modeling, defense mechanisms, and evaluation validation. Through comprehensive literature analysis and comparative study, we categorize adversarial attack vectors—including observation perturbations and instruction injection—and survey robustness-oriented defenses, such as perception-planning co-hardening and safety-constraint embedding. We further identify critical limitations in current evaluations, including dataset bias and metric oversimplification, and uncover key challenges: expanding attack surfaces, cross-modal vulnerabilities, and lack of trustworthy verification. Based on this analysis, we propose four future research directions: (1) interpretable attack taxonomies; (2) hierarchical defense frameworks; (3) multi-granularity safety evaluation benchmarks; and (4) formal verification toolchains. These contributions provide both theoretical foundations and practical guidelines for developing secure and reliable embodied navigation systems.
This study addresses the limitations of current robotic system validation, which relies heavily on manual selection of test scenarios, thereby hindering scalability and compromising reproducibility and reliability of conclusions. To overcome these challenges, this work proposes a compositional, scenario-based modeling approach that integrates declarative test specifications, plugin-driven scenario generation, containerized parallel simulation, and unified result analysis to establish the first modular and scalable automated verification framework. The framework enables systematic parameter variation across multiple dimensions and facilitates robust identification of systemic faults versus stochastic anomalies. Evaluated across 5,480 distinct scenario configurations with over 100,000 simulation runs, the approach accumulated 1,800 hours of simulated operation and 1,873 virtual kilometers, demonstrating its efficacy in discerning consistent system deficiencies from random irregularities.
This work addresses the performance degradation commonly observed in sim-to-real transfer of reinforcement learning for robotic navigation, which stems from domain discrepancies and a lack of systematic analysis linking training strategies to deployment outcomes. The authors propose an end-to-end training and deployment pipeline that decouples key influencing factors, introducing perturbation-aware fine-tuning and a Transformer-based temporal reasoning policy to significantly enhance zero-shot transfer robustness and control smoothness. By integrating perturbation modeling, post-training fine-tuning, and system-level domain gap analysis, the method outperforms existing learning-based baselines in both static and dynamic environments, matching the performance of optimization-based planners in static scenes and achieving successful zero-shot deployment across multiple real-world robotic platforms.
In off-road autonomous navigation, motion uncertainty arising from unmanned ground vehicle (UGV)–terrain interaction cannot be directly observed by onboard sensors, limiting motion model fidelity. To address this, we propose DRIVE, a slip system identification protocol that enables standardized data collection under steady-state slip conditions—covering six terrain types, multiple platforms (75–470 kg), and 14.7 km of real-world testing. We develop a transfer function model linking commanded velocity to steady-state slip. Furthermore, we introduce the first terrain–robot interaction risk metric—“command uncertainty”—defined based on steady-state response, enabling probabilistic and interpretable quantification of slip-induced navigation risk. Experimental validation demonstrates that this metric accurately identifies high-risk slip scenarios, providing reliable, risk-aware inputs for autonomous navigation decision-making.
Industrial quadruped robots face challenges in robustly navigating dynamic environments, and conventional testing methods suffer from low coverage and poor reproducibility. Method: This paper pioneers the adaptation of Surrealist—a search-based simulation testing framework originally developed for UAVs—to the ANYmal quadruped platform. We propose an automated, closed-loop scenario generation and verification methodology integrating high-fidelity simulation modeling, evolutionary scene mutation strategies, and quantitative success-rate evaluation. The approach enables objective, black-box comparison and systematic defect exposure for proprietary navigation algorithms. Contribution/Results: In pilot deployment, our framework identified a critical performance bottleneck—40.3% success rate—for one navigation algorithm, while verifying another achieving 71.2%. Within six months, it enabled efficient, repeatable evaluation of five distinct algorithms. The method significantly enhances automation, reproducibility, and rigor in industrial-grade navigation system validation.
Learning-based black-box autonomous mobile robots struggle to satisfy dynamically evolving human safety requirements. Method: This paper proposes a regulator-driven, post-hoc safety assessment framework. Its core innovations include: (i) systematically modeling human safety requirements as Signal Temporal Logic (STL) specifications for the first time; (ii) introducing differentiable, quantitative safety metrics—Total Robustness Value (TRV) and Local Robustness Value (LRV); and (iii) enabling closed-loop model retraining via external trajectory verification and robustness feedback. Results: In virtual driving tasks, speeding and lane-deviation violations decreased by 177% and 1138%, respectively. In robot navigation experiments, sharp-turn evasive capability improved by 300%, and time-to-collision with obstacles reduced by 49%. Real-world robotic deployment validates both effectiveness and generalizability.
Sim-to-real transfer in robotics is fundamentally hindered by the *reality gap*—systematic discrepancies between simulation and reality in dynamics, perception, and interaction. This work systematically analyzes the root causes of the reality gap and proposes a unified, causality-driven conceptual framework. We comprehensively survey and comparatively evaluate mainstream mitigation strategies, including domain randomization, sim-real co-training, state-action abstraction, and real-data feedback. Drawing on empirical insights, we distill cross-platform transfer best practices. Our key contribution is a closed-loop analytical framework integrating *causes*, *methods*, and *evaluation*, validated across diverse robotic tasks—including navigation, locomotion control, and dexterous manipulation. Experimental results demonstrate significant improvements in generalization performance and a measurable reduction in sim-to-real performance degradation, thereby advancing the practical deployment of learned robotic policies.
This study addresses the degradation of control performance and inefficient resource utilization in cooperative robotic navigation under complex environments, where wireless latency and fluctuating communication reliability significantly impair system efficacy. For the first time, a Quality of Control (QoC) framework is extended to real-world multi-robot navigation systems, employing closed-loop control modeling to quantitatively assess the impact of network effects on control performance and systematically analyze the coupling between control parameters and communication Quality of Service (QoS). Leveraging a private 5G testbed and empirical evaluations of diverse ROS 2 QoS policies, the work identifies an operational regime for joint control-communication optimization. Experimental results demonstrate that, in representative scenarios, the RELIABLE QoS policy improves QoC by up to 51.5% compared to BEST_EFFORT, offering a principled basis for optimal cooperative configuration in practical robotic systems.
This work addresses the limited reproducibility of behavioral validation in robotic simulation testing, which often stems from insufficiently documented test configurations, execution protocols, and post-processing procedures. To overcome this, the study proposes a deep integration of data provenance and FAIR (Findable, Accessible, Interoperable, Reusable) principles throughout the entire test generation pipeline—rather than merely appending them to final datasets. The authors extend an existing simulation testing framework by embedding machine-readable, structured metadata at every stage, thereby enabling end-to-end traceable validation workflows. This approach significantly enhances the reproducibility of mobile robot navigation datasets. Additionally, the project distills practical FAIR implementation guidelines tailored to robotics, identifying key challenges such as vocabulary alignment, attribute selection, and adoption of community standards, and offers actionable recommendations for addressing them.
This work addresses the lack of reproducible benchmarks for systematically evaluating zero-shot sim-to-real transfer of multi-agent reinforcement learning (MARL) policies in connected autonomous driving. To bridge this gap, the authors establish a unified benchmark that integrates high-fidelity digital twins, a physical testbed, and simulation environments within the Cyber-Physical Mobility Lab framework. For the first time, this setup enables end-to-end, zero-shot deployment and evaluation of MARL policies across all three domains under rigorously reproducible real-world conditions. By deploying the SigmaRL policy, the study quantitatively reveals how discrepancies in control architectures and environmental fidelity critically contribute to performance degradation during transfer. The resulting framework provides an open-source, structured, and reproducible foundation for advancing MARL research in sim-to-real transfer.
This work addresses the weak correlation between conventional offline evaluation metrics and actual robotic policy performance, which hinders efficient policy iteration. To bridge this gap, the authors propose Critical Interval Mean Squared Error (CI-MSE), a novel offline evaluation method that focuses error assessment on task-critical time intervals and incorporates a lightweight action alignment mechanism to better reflect real-world deployment outcomes. Evaluated across multiple policy checkpoints, CI-MSE achieves a Spearman correlation coefficient of −0.87 with true roll-out performance—significantly outperforming standard MSE (−0.61)—and demonstrates strong robustness to hyperparameter variations. These results indicate that CI-MSE substantially enhances the reliability and practical utility of offline policy validation in robotics.