navigation evaluation metrics

Defining and applying evaluation metrics and validation protocols that measure robustness and generalization of navigation systems across simulation and real-world robot deployments, ensuring improvements transfer between environments.

navigationevaluationmetrics

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Safety of Embodied Navigation: A Survey

Aug 07, 2025
ZW
Zixia Wang
🏛️ University of Exeter

This paper addresses safety risks confronting embodied navigation systems in dynamic real-world environments. We present the first systematic survey spanning three dimensions: attack modeling, defense mechanisms, and evaluation validation. Through comprehensive literature analysis and comparative study, we categorize adversarial attack vectors—including observation perturbations and instruction injection—and survey robustness-oriented defenses, such as perception-planning co-hardening and safety-constraint embedding. We further identify critical limitations in current evaluations, including dataset bias and metric oversimplification, and uncover key challenges: expanding attack surfaces, cross-modal vulnerabilities, and lack of trustworthy verification. Based on this analysis, we propose four future research directions: (1) interpretable attack taxonomies; (2) hierarchical defense frameworks; (3) multi-granularity safety evaluation benchmarks; and (4) formal verification toolchains. These contributions provide both theoretical foundations and practical guidelines for developing secure and reliable embodied navigation systems.

Analyzing safety concerns in embodied navigation systemsExploring attack and defense strategies for navigation safetyIdentifying future research directions for reliable navigation

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limitations of current robotic system validation, which relies heavily on manual selection of test scenarios, thereby hindering scalability and compromising reproducibility and reliability of conclusions. To overcome these challenges, this work proposes a compositional, scenario-based modeling approach that integrates declarative test specifications, plugin-driven scenario generation, containerized parallel simulation, and unified result analysis to establish the first modular and scalable automated verification framework. The framework enables systematic parameter variation across multiple dimensions and facilitates robust identification of systemic faults versus stochastic anomalies. Evaluated across 5,480 distinct scenario configurations with over 100,000 simulation runs, the approach accumulated 1,800 hours of simulated operation and 1,873 virtual kilometers, demonstrating its efficacy in discerning consistent system deficiencies from random irregularities.

automated testingreproducibilityrobot validation

This work addresses the performance degradation commonly observed in sim-to-real transfer of reinforcement learning for robotic navigation, which stems from domain discrepancies and a lack of systematic analysis linking training strategies to deployment outcomes. The authors propose an end-to-end training and deployment pipeline that decouples key influencing factors, introducing perturbation-aware fine-tuning and a Transformer-based temporal reasoning policy to significantly enhance zero-shot transfer robustness and control smoothness. By integrating perturbation modeling, post-training fine-tuning, and system-level domain gap analysis, the method outperforms existing learning-based baselines in both static and dynamic environments, matching the performance of optimization-based planners in static scenes and achieving successful zero-shot deployment across multiple real-world robotic platforms.

domain discrepancyreal-world deploymentreinforcement learning

In off-road autonomous navigation, motion uncertainty arising from unmanned ground vehicle (UGV)–terrain interaction cannot be directly observed by onboard sensors, limiting motion model fidelity. To address this, we propose DRIVE, a slip system identification protocol that enables standardized data collection under steady-state slip conditions—covering six terrain types, multiple platforms (75–470 kg), and 14.7 km of real-world testing. We develop a transfer function model linking commanded velocity to steady-state slip. Furthermore, we introduce the first terrain–robot interaction risk metric—“command uncertainty”—defined based on steady-state response, enabling probabilistic and interpretable quantification of slip-induced navigation risk. Experimental validation demonstrates that this metric accurately identifies high-risk slip scenarios, providing reliable, risk-aware inputs for autonomous navigation decision-making.

Evaluating terrain-robot interaction reachable velocities using DRIVE protocolProposing unpredictability metric for command uncertainty and risk assessmentStandardizing data collection for off-road autonomous navigation slip states

Bridging Research and Practice in Simulation-based Testing of Industrial Robot Navigation Systems

Oct 10, 2025
SK
Sajad Khatiri
🏛️ Università della Svizzera italiana | University of Bern | ANYbotics AG

Industrial quadruped robots face challenges in robustly navigating dynamic environments, and conventional testing methods suffer from low coverage and poor reproducibility. Method: This paper pioneers the adaptation of Surrealist—a search-based simulation testing framework originally developed for UAVs—to the ANYmal quadruped platform. We propose an automated, closed-loop scenario generation and verification methodology integrating high-fidelity simulation modeling, evolutionary scene mutation strategies, and quantitative success-rate evaluation. The approach enables objective, black-box comparison and systematic defect exposure for proprietary navigation algorithms. Contribution/Results: In pilot deployment, our framework identified a critical performance bottleneck—40.3% success rate—for one navigation algorithm, while verifying another achieving 71.2%. Within six months, it enabled efficient, repeatable evaluation of five distinct algorithms. The method significantly enhances automation, reproducibility, and rigor in industrial-grade navigation system validation.

Automating obstacle scenario generation for quadrupedal robot inspectionTesting robotic navigation robustness in dynamic industrial environmentsValidating algorithm performance through simulation-based failure detection

SAFE-SMART: Safety Analysis and Formal Evaluation using STL Metrics for Autonomous RoboTs

Nov 21, 2025
KS
Kristy Sakano
🏛️ University of Maryland, College Park

Learning-based black-box autonomous mobile robots struggle to satisfy dynamically evolving human safety requirements. Method: This paper proposes a regulator-driven, post-hoc safety assessment framework. Its core innovations include: (i) systematically modeling human safety requirements as Signal Temporal Logic (STL) specifications for the first time; (ii) introducing differentiable, quantitative safety metrics—Total Robustness Value (TRV) and Local Robustness Value (LRV); and (iii) enabling closed-loop model retraining via external trajectory verification and robustness feedback. Results: In virtual driving tasks, speeding and lane-deviation violations decreased by 177% and 1138%, respectively. In robot navigation experiments, sharp-turn evasive capability improved by 300%, and time-to-collision with obstacles reduced by 49%. Real-world robotic deployment validates both effectiveness and generalizability.

Ensuring compliance with human-defined safety rules using STL specificationsPost hoc safety evaluation of black-box autonomous mobile robotsQuantitative safety metrics for iterative improvement of robot behavior

Latest Papers

What's happening recently
View more

The Reality Gap in Robotics: Challenges, Solutions, and Best Practices

Oct 23, 2025
EA
Elie Aljalbout
🏛️ University of Zurich | NVIDIA | University of Washington | University of Utah | The University of Sydney

Sim-to-real transfer in robotics is fundamentally hindered by the *reality gap*—systematic discrepancies between simulation and reality in dynamics, perception, and interaction. This work systematically analyzes the root causes of the reality gap and proposes a unified, causality-driven conceptual framework. We comprehensively survey and comparatively evaluate mainstream mitigation strategies, including domain randomization, sim-real co-training, state-action abstraction, and real-data feedback. Drawing on empirical insights, we distill cross-platform transfer best practices. Our key contribution is a closed-loop analytical framework integrating *causes*, *methods*, and *evaluation*, validated across diverse robotic tasks—including navigation, locomotion control, and dexterous manipulation. Experimental results demonstrate significant improvements in generalization performance and a measurable reduction in sim-to-real performance degradation, thereby advancing the practical deployment of learned robotic policies.

Addressing discrepancies between simulated and real robotic environmentsOvercoming challenges in transferring systems from simulation to realityProviding comprehensive overview of reality gap causes and solutions

This study addresses the degradation of control performance and inefficient resource utilization in cooperative robotic navigation under complex environments, where wireless latency and fluctuating communication reliability significantly impair system efficacy. For the first time, a Quality of Control (QoC) framework is extended to real-world multi-robot navigation systems, employing closed-loop control modeling to quantitatively assess the impact of network effects on control performance and systematically analyze the coupling between control parameters and communication Quality of Service (QoS). Leveraging a private 5G testbed and empirical evaluations of diverse ROS 2 QoS policies, the work identifies an operational regime for joint control-communication optimization. Experimental results demonstrate that, in representative scenarios, the RELIABLE QoS policy improves QoC by up to 51.5% compared to BEST_EFFORT, offering a principled basis for optimal cooperative configuration in practical robotic systems.

collaborative navigationquality of controlreliability variations

This work addresses the limited reproducibility of behavioral validation in robotic simulation testing, which often stems from insufficiently documented test configurations, execution protocols, and post-processing procedures. To overcome this, the study proposes a deep integration of data provenance and FAIR (Findable, Accessible, Interoperable, Reusable) principles throughout the entire test generation pipeline—rather than merely appending them to final datasets. The authors extend an existing simulation testing framework by embedding machine-readable, structured metadata at every stage, thereby enabling end-to-end traceable validation workflows. This approach significantly enhances the reproducibility of mobile robot navigation datasets. Additionally, the project distills practical FAIR implementation guidelines tailored to robotics, identifying key challenges such as vocabulary alignment, attribute selection, and adoption of community standards, and offers actionable recommendations for addressing them.

data provenanceFAIR principlesreplicability

This work addresses the lack of reproducible benchmarks for systematically evaluating zero-shot sim-to-real transfer of multi-agent reinforcement learning (MARL) policies in connected autonomous driving. To bridge this gap, the authors establish a unified benchmark that integrates high-fidelity digital twins, a physical testbed, and simulation environments within the Cyber-Physical Mobility Lab framework. For the first time, this setup enables end-to-end, zero-shot deployment and evaluation of MARL policies across all three domains under rigorously reproducible real-world conditions. By deploying the SigmaRL policy, the study quantitatively reveals how discrepancies in control architectures and environmental fidelity critically contribute to performance degradation during transfer. The resulting framework provides an open-source, structured, and reproducible foundation for advancing MARL research in sim-to-real transfer.

BenchmarkConnected and Automated VehiclesMulti-Agent Reinforcement Learning

This work addresses the weak correlation between conventional offline evaluation metrics and actual robotic policy performance, which hinders efficient policy iteration. To bridge this gap, the authors propose Critical Interval Mean Squared Error (CI-MSE), a novel offline evaluation method that focuses error assessment on task-critical time intervals and incorporates a lightweight action alignment mechanism to better reflect real-world deployment outcomes. Evaluated across multiple policy checkpoints, CI-MSE achieves a Spearman correlation coefficient of −0.87 with true roll-out performance—significantly outperforming standard MSE (−0.61)—and demonstrates strong robustness to hyperparameter variations. These results indicate that CI-MSE substantially enhances the reliability and practical utility of offline policy validation in robotics.

offline validationperformance correlationpolicy evaluation

Hot Scholars

IK

Itzik Klein

University of Haifa
RoboticsInertial SensingData-Driven NavigationAUV
DS

Daeun Song

Postdoctoral Associate, George Mason University
Robot Path and Motion PlanningComputational Geometry
DM

Dinesh Manocha

Distinguished University Professor, University of Maryland at College Park
computer graphicsgeometric modelingmotion planningvirtual reality
XH

Xiaoshuai Hao

Beijing Academy of Artificial Intelligence,BAAI
vision and language