๐ค AI Summary
This work addresses robot navigation in environments characterized by spatially correlated obstacles, uncertain obstacle states, high-sensor noise, and costly obstacle identification. To tackle these challenges, we propose a Gaussian random field-based Bayesian belief update mechanism that efficiently models the joint posterior distribution over obstacles. We design a two-stage learning framework: (i) an information-theoretic reward-driven exploratory perception phase, and (ii) a joint optimization of expected path cost and risk-sensitive policies via Monte Carlo point estimation and distributional reinforcement learning. Our key contributions include correlation-aware belief updating and optimistic policy iteration, enabling full-cost-distribution modeling and fine-grained uncertainty quantification. Experiments demonstrate that our method consistently outperforms baselines across varying obstacle densities and sensor accuracies, and exhibits strong robustness and adaptability in complex, adversarial settingsโsuch as those involving interference or clustered hazards.
๐ Abstract
We introduce the Stochastic Correlated Obstacle Scene (SCOS) problem, a navigation setting with spatially correlated obstacles of uncertain blockage status, realistically constrained sensors that provide noisy readings and costly disambiguation. Modeling the spatial correlation with Gaussian Random Field (GRF), we develop Bayesian belief updates that refine blockage probabilities, and use the posteriors to reduce search space for efficiency. To find the optimal traversal policy, we propose a novel two-stage learning framework. An offline phase learns a robust base policy via optimistic policy iteration augmented with information bonus to encourage exploration in informative regions, followed by an online rollout policy with periodic base updates via a Bayesian mechanism for information adaptation. This framework supports both Monte Carlo point estimation and distributional reinforcement learning (RL) to learn full cost distributions, leading to stronger uncertainty quantification. We establish theoretical benefits of correlation-aware updating and convergence property under posterior sampling. Comprehensive empirical evaluations across varying obstacle densities, sensor capabilities demonstrate consistent performance gains over baselines. This framework addresses navigation challenges in environments with adversarial interruptions or clustered natural hazards.