FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of predicting floor affordance (FloAff) for mobile manipulation, where success depends not only on navigational feasibility but also on task-specific manipulability. To this end, the authors propose a unified framework that leverages egocentric multimodal perception to enable robust FloAff prediction. The key innovations include the introduction of a canonical floor affordance representation (CFAR) to eliminate irrelevant spatial variations, and a progressive FloAff learning (PFAL) strategy that effectively transfers priors from foundational tasks to heterogeneous downstream manipulation skills. The study further presents FloAff-Kitchen, the first cross-scene, multi-view benchmark for floor affordance, and demonstrates significant performance gains over strong baselines across three evaluation settings. Ablation studies confirm the contribution of each component, and both code and dataset are publicly released.
📝 Abstract
Mobile manipulation requires robots to identify Floor Affordance (FloAff) that maximizes downstream manipulation success rather than merely ensuring navigation feasibility. FloAff prediction is a target-conditioned local spatial reasoning problem, yet existing methods suffer from representation ambiguity caused by irrelevant spatial context and arbitrary object orientations, while entangling shared and task-specific knowledge across heterogeneous manipulation skills. To address these challenges, we propose a unified framework for FloAff prediction from egocentric multimodal perception, consisting of canonical representation learning and progressive affordance prior learning. Specifically, we introduce a Canonical Floor Affordance Representation (CFAR), which learns canonical interaction geometry by preserving affordance-relevant local structure while eliminating nuisance spatial variations unrelated to robot base placement. We further propose Progressive Floor Affordance Learning (PFAL), which learns transferable FloAff priors from a foundation manipulation task and progressively adapts them to heterogeneous downstream manipulation skills. To facilitate systematic evaluation, we establish the first cross-scene, multi-view FloAff-Kitchen benchmark covering diverse manipulation skills, scene layouts, furniture styles, and viewpoints. Extensive experiments on three benchmark settings demonstrate that our method consistently outperforms strong baselines, while ablation studies validate the contribution of each proposed component. Project page: https://csu-hero-lab.github.io/FloAff-Kitchen_Web/
Problem

Research questions and friction points this paper is trying to address.

Floor Affordance
mobile manipulation
spatial reasoning
representation ambiguity
heterogeneous manipulation skills
Innovation

Methods, ideas, or system contributions that make the work stand out.

Floor Affordance
Canonical Representation
Progressive Learning
Mobile Manipulation
Multimodal Perception