PALMS+: Modular Image-Based Floor Plan Localization Leveraging Depth Foundation Model

📅 2025-11-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of infrastructure-free indoor localization under GPS-denied conditions (e.g., emergency response and assistive navigation), this paper proposes PALMS+, a monocular visual localization framework. Methodologically, it reconstructs metric-scale 3D point clouds from single RGB frames using the Depth Pro model, then performs convolutional matching against floorplan geometry and posterior probability inference—enabling high-accuracy static and sequential localization without any training, and supporting particle-filter-based trajectory tracking. Key contributions include: (i) the first integration of monocular depth estimation with floorplan-aware geometric convolutional matching, circumventing LiDAR’s range limitations and floorplan layout ambiguities; and (ii) zero-shot, metric-scale-consistent, modular localization. Evaluated on Structured3D and real-world campus scenes, PALMS+ achieves superior static localization accuracy across 80 observation points compared to PALMS and F3Loc, and attains lower average trajectory error over 33 sequences, demonstrating robustness and practical applicability.

Technology Category

Intelligent Robots: Localization, Mapping, and NavigationComputer Vision: Motion & TrackingPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsSecurity and Privacy: Large-scale security measurementsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Indoor localization in GPS-denied environments is crucial for applications like emergency response and assistive navigation. Vision-based methods such as PALMS enable infrastructure-free localization using only a floor plan and a stationary scan, but are limited by the short range of smartphone LiDAR and ambiguity in indoor layouts. We propose PALMS$+$, a modular, image-based system that addresses these challenges by reconstructing scale-aligned 3D point clouds from posed RGB images using a foundation monocular depth estimation model (Depth Pro), followed by geometric layout matching via convolution with the floor plan. PALMS$+$ outputs a posterior over the location and orientation, usable for direct or sequential localization. Evaluated on the Structured3D and a custom campus dataset consisting of 80 observations across four large campus buildings, PALMS$+$ outperforms PALMS and F3Loc in stationary localization accuracy -- without requiring any training. Furthermore, when integrated with a particle filter for sequential localization on 33 real-world trajectories, PALMS$+$ achieved lower localization errors compared to other methods, demonstrating robustness for camera-free tracking and its potential for infrastructure-free applications. Code and data are available at https://github.com/Head-inthe-Cloud/PALMS-Plane-based-Accessible-Indoor-Localization-Using-Mobile-Smartphones
Problem

Research questions and friction points this paper is trying to address.

Addresses indoor localization in GPS-denied environments using visual data
Overcomes limitations of smartphone LiDAR range and layout ambiguity
Provides infrastructure-free localization through modular image-based reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modular image-based system reconstructs scale-aligned 3D point clouds
Uses foundation monocular depth estimation model for reconstruction
Performs geometric layout matching via convolution with floor plan
🔎 Similar Papers
2024-03-05Computer Vision and Pattern RecognitionCitations: 4
💼 Related Jobs
No related jobs found.
Y
Yunqian Cheng
University of California, Santa Cruz
B
Benjamin Princen
University of California, Santa Cruz
Roberto Manduchi
Roberto Manduchi
Professor of Computer Science and Engineering, UC Santa Cruz
Computer visionAssistive TechnologyImage ProcessingSensors