🤖 AI Summary
This study addresses the challenges of sparse rewards and failure attribution arising from long-horizon decision-making in multi-floor object navigation by proposing a modular diagnostic framework. Methodologically, a hierarchical factorization strategy decouples the global policy into intra-floor exploration and inter-floor transition sub-policies, with theoretical proof establishing their equivalence to the original global policy. Additionally, knowledge distillation from vision-language models is introduced to enable efficient reinforcement learning initialization. Experiments reveal that perception performance and stair-climbing stability constitute the core bottlenecks limiting navigation success rates. By establishing a standardized diagnostic paradigm for multi-floor navigation failures, this work offers a novel, interpretable pathway for policy decomposition in embodied navigation within complex environments.
📝 Abstract
Object-goal navigation (ObjectNav) in multi-floor scenarios presents a challenge due to sparse rewards caused by long-horizon decision-making. In this paper, we propose a diagnostic study based on a modular framework with an effective learnable policy to analyze failure factors in multi-floor scenarios. To achieve an effective policy for diagnosis, we design the hierarchical factorization policy that deconstructs a single global policy into an intra-floor exploration policy and an inter-floor switching policy. To providing an effective initialization for Reinforcement Learning (RL), the lightweight intra-floor policy is learned by distilling the exploration logic of Visual Language Models (VLMs). Under idealized assumptions, we show that the factorized policy is theoretically equivalent to a single global policy at the policy-representation level. Experiment results indicate that perception performance and stair climbing stability are the primary bottlenecks in multi-floor navigation.