Score
Design, build, or evaluate models and algorithms that detect and localize parts, surfaces, or specific points on objects that enable particular interactions, producing affordance labels or interaction locations and associating them with actionable verbs or tasks; includes methods that generalize affordance prediction to novel categories or open-vocabulary descriptions.
Current visual affordance prediction research suffers from fragmented task definitions—such as grasp detection and affordance classification—each redefining “affordance” independently, leading to incomparable benchmarks and poor reproducibility, thereby hindering generalization in robotic interaction. To address this, we propose a unified modeling framework: (1) formally define visual affordance as a mapping from object visual features to physical attributes (e.g., mass) and subsequently to task-relevant interactions; (2) introduce the *Affordance Sheet*, a standardized documentation protocol specifying data curation, evaluation metrics, and implementation details; and (3) establish a cross-task benchmark and reproducibility diagnostic toolkit integrating visual perception, physical attribute estimation, and interaction modeling. This work is the first to achieve conceptual unification, evaluation consistency, and implementation transparency in affordance prediction, significantly improving model reliability and generalization in real-world robotic settings.
This work addresses open-vocabulary, task-oriented robotic grasping in cluttered scenes. We propose the first vision-language model (VLM)-based framework for implicit instruction-driven affordance reasoning—requiring no explicit task annotations. Our method jointly leverages in-context prompting, CLIP/LLaVA for semantic understanding, a learnable visual grounding module, and a geometry-aware grasp pose generation network to implicitly infer task goals from natural language instructions, localize relevant objects, and output functionally consistent part-level grasp poses. Key contributions include zero-shot generalization to unseen objects and open-vocabulary tasks, eliminating reliance on supervised data of fixed task–object pairs. Evaluated in both simulation and real-world settings, our approach achieves state-of-the-art performance, improving task success rate by 27.3% over baselines and successfully generalizing to over 100 novel objects and 50+ open-ended instructions.
To address the poor generalization and deployment challenges of object affordance reasoning in task-oriented robotic manipulation, this paper proposes an end-to-end vision-action semantic mapping framework. Methodologically, we introduce LVIS-Aff—a large-scale, multi-task affordance dataset—and design Afford-X, a lightweight model featuring novel Verb Attention and Bidirectional Cross-Modal Fusion (Bi-Fusion) modules to enable perception-driven affordance modeling and efficient edge inference. Contributions include: (1) a 12.1% absolute performance gain over prior non-LLM approaches (+1.2% relative improvement), (2) only 187M parameters, and (3) inference speed 50× faster than the GPT-4V API. The framework is validated across multiple robotic platforms and real-world environments, demonstrating strong generalizability and practical deployability.
Open-vocabulary 3D object affordance localization aims to precisely localize functional regions on 3D objects that enable actions specified by arbitrary natural language instructions. Existing methods suffer from limited semantic priors and struggle to model geometric invariance and implicit interaction intent. To address this, we propose a geometry-intent collaborative reasoning paradigm: (i) an implicit invariant geometric representation that disentangles intrinsic structural properties; (ii) an interaction intent analogy mechanism enabling human-like functional reasoning across objects and instructions; and (iii) PIADv2—the largest open-vocabulary 3D affordance understanding dataset to date. Our method integrates point-cloud–image joint representation learning, geometry-guided intent distillation, and contrastive analogy reasoning. It achieves significant improvements over state-of-the-art methods on open-vocabulary affordance localization, demonstrating strong generalization to unseen objects and instructions. Code and dataset are publicly released.
This work addresses the limited generalization of existing robotic planning systems, which rely heavily on visual appearance and neglect task-relevant functional properties, thereby struggling in novel robot-object interaction scenarios. To overcome this, the authors propose A4D, a novel approach that establishes a shared latent space centered on functional attributes (e.g., “movable”). A4D enables efficient reasoning through alignment of visual and functional embeddings and proximity-driven matching to functional prototypes. Furthermore, it incorporates an uncertainty-based few-shot mechanism for discovering new functions in previously unseen contexts. The method achieves 94% accuracy in known-function inference—surpassing prior approaches by over 15 percentage points—and, for novel functions, requires less than 10% of the training data to boost accuracy from 70% to above 90%, while also delivering a hundredfold speedup in inference time.
Existing approaches struggle to accurately localize multifunctional regions in cluttered real-world scenes according to task instructions and lack benchmarks supporting complex mappings such as one-to-many task-to-region correspondences. This work proposes a task-conditioned, scene-level functional affordance localization framework and introduces the first real-scene affordance benchmark dataset encompassing both single- and multi-region instruction mappings. To enable efficient and high-quality annotation, we design the A2A-AffordGen pipeline, which integrates large language model filtering, interactive segmentation, mask refinement, and human verification. Models trained on this dataset significantly outperform baseline methods—including general-purpose segmentation models, vision-language models, and affordance distillation approaches—in both task-level localization accuracy and spatial priors for downstream manipulation tasks.
This work proposes the first open-world general-purpose affordance foundation model, unifying the core challenges of "where to interact" and "how to interact." Given only a single RGB-D image and a language instruction, the model predicts task-relevant functional region masks and 3D post-contact motion trajectories. A large-scale, standardized data pipeline integrates robotic manipulation, human demonstrations, simulation, and real-world scan data to construct a unified language–mask–3D motion affordance representation. Evaluated across eight benchmarks, the model substantially outperforms existing methods, achieving average gains of 23.9 in gIoU and 26.3 in cIoU, improving contact-point hit rates by 12.7–61.3%, and demonstrating superior 3D motion prediction. Notably, it enables zero-shot generalization across objects, tasks, and scenes and can be deployed directly on real robots without fine-tuning.
This work addresses the problem of affordance grounding for robotic tool use in open-world settings—specifically, selecting appropriate tools from arbitrary object categories and precisely localizing their functional regions. The authors propose a hierarchical grounding framework that treats object parts as abstract units. This approach leverages vision-language models for task parsing, tool selection, and part identification, and integrates foundation vision models to accurately map identified parts to 3D manipulation regions, all from a single RGB-D image. Without requiring large-scale end-to-end training, the framework outperforms existing methods on standard affordance prediction benchmarks and demonstrates zero-shot generalization to open-category tools in both simulated and real-world robotic experiments.
This work addresses the challenge of enabling robots to reliably plan and execute tasks from natural language instructions in complex, dynamic environments with occlusions, while achieving effective sim-to-real transfer. To this end, the authors propose a planning framework that integrates functional affordance recognition with visual action-effect prediction, leveraging visual forward reasoning to anticipate future states. A multimodal text-image matching module is introduced to evaluate the consistency between candidate action sequences and the linguistic goal. Furthermore, a real-to-sim image stylization mechanism is designed to enhance perceptual robustness in real-world settings. Experimental results demonstrate that the proposed approach successfully accomplishes challenging manipulation tasks on both simulated and physical robot platforms, significantly improving language-conditioned generalization from simulation to reality.