From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of multimodal pose distribution in visual floorplan localization, which arises from cross-modal information asymmetry and repetitive indoor layouts. To this end, we propose a coarse-to-fine localization framework that eliminates the need for ray matching. Our approach introduces, for the first time, an image-conditioned pose diffusion model to capture continuous multimodal pose distributions, complemented by a local refinement module that achieves sub-meter accuracy within candidate regions. By employing candidate-centered floorplan cropping and enabling end-to-end inference without offline preprocessing, the system attains state-of-the-art accuracy and robustness on the S3D (full) and ZInD benchmarks, departing from conventional paradigms that rely on precomputed maps or lookup tables during testing.
📝 Abstract
Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.
Problem

Research questions and friction points this paper is trying to address.

Visual Floorplan Localization
multimodal pose distribution
indoor localization
cross-modal asymmetry
structural ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

pose diffusion
coarse-to-fine localization
visual floorplan localization
multimodal pose distribution
ray-free matching
🔎 Similar Papers
2024-03-05Computer Vision and Pattern RecognitionCitations: 4