🤖 AI Summary
Current visual affordance prediction research suffers from fragmented task definitions—such as grasp detection and affordance classification—each redefining “affordance” independently, leading to incomparable benchmarks and poor reproducibility, thereby hindering generalization in robotic interaction. To address this, we propose a unified modeling framework: (1) formally define visual affordance as a mapping from object visual features to physical attributes (e.g., mass) and subsequently to task-relevant interactions; (2) introduce the *Affordance Sheet*, a standardized documentation protocol specifying data curation, evaluation metrics, and implementation details; and (3) establish a cross-task benchmark and reproducibility diagnostic toolkit integrating visual perception, physical attribute estimation, and interaction modeling. This work is the first to achieve conceptual unification, evaluation consistency, and implementation transparency in affordance prediction, significantly improving model reliability and generalization in real-world robotic settings.
📝 Abstract
Human-robot interaction for assistive technologies relies on the prediction of affordances, which are the potential actions a robot can perform on objects. Predicting object affordances from visual perception is formulated differently for tasks such as grasping detection, affordance classification, affordance segmentation, and hand-object interaction synthesis. In this work, we highlight the reproducibility issue in these redefinitions, making comparative benchmarks unfair and unreliable. To address this problem, we propose a unified formulation for visual affordance prediction, provide a comprehensive and systematic review of previous works highlighting strengths and limitations of methods and datasets, and analyse what challenges reproducibility. To favour transparency, we introduce the Affordance Sheet, a document to detail the proposed solution, the datasets, and the validation. As the physical properties of an object influence the interaction with the robot, we present a generic framework that links visual affordance prediction to the physical world. Using the weight of an object as an example for this framework, we discuss how estimating object mass can affect the affordance prediction. Our approach bridges the gap between affordance perception and robot actuation, and accounts for the complete information about objects of interest and how the robot interacts with them to accomplish its task.