affordance-aware rag

Design and build retrieval-augmented generation systems that represent, index, and retrieve object and action affordances—including hierarchical or multi-dimensional affordance descriptors—so retrieval and ranking are driven by functional compatibility rather than visual similarity. Implement affordance encoding, compatibility-based retrieval and prioritization, and generation components that produce affordance-grounded strategies and grasp/action suggestions.

affordance-awarerag

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing affordance prediction methods struggle in unseen environments due to sparse retrieval, limited generalization, and inaccurate contact point localization. This work proposes the RAAP framework, which innovatively integrates retrieval augmentation with cross-image action alignment. By decoupling static contact point localization from dynamic action direction prediction, RAAP leverages dense image correspondences to transfer contact points and introduces a dual-weighted attention mechanism to fuse multiple reference samples for robust action direction estimation. Requiring only minimal training data, the method achieves strong affordance prediction on novel objects and categories. Trained on small-scale subsets of DROID and HOI4D, RAAP successfully enables zero-shot robotic manipulation in both simulation and real-world settings, significantly enhancing cross-category generalization and robustness.

affordance predictioncontact localizationgeneralization

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

Dec 21, 2025
RK
Ryosuke Korekata
🏛️ Keio University | Carnegie Mellon University

This work addresses open-vocabulary mobile manipulation tasks, where robots must accurately localize and manipulate diverse objects in unseen indoor environments guided by free-form natural language instructions. We propose a hierarchical multimodal retrieval framework coupled with affordance-aware embodied memory modeling: for the first time, we decouple visual region semantics from physical affordances to construct an Affordance-Aware Embodied Memory; further integrating vision-language models, region-level image embeddings, a functional scoring network, and a hierarchical retrieval architecture to enable zero-shot hierarchical retrieval and re-ranking. Our method achieves significantly superior retrieval performance over state-of-the-art approaches on large-scale benchmarks; in real-robot experiments, it attains an 85% task success rate—marking the first demonstration of highly robust, instruction-driven mobile manipulation under open-vocabulary conditions.

Open-vocabulary mobile manipulation with natural language instructionsRetrieving and ranking executable manipulation options for robotsUnderstanding visual semantics and manipulation affordances in real environments

This work addresses the limitations of current vision-language model (VLM)-driven robotic grasping methods, which rely heavily on visual similarity while neglecting physical affordances—such as graspable regions and material fragility—and lack spatial reasoning and failure recovery capabilities, leading to poor generalization in dense, cluttered scenes. To overcome these challenges, the authors propose an agent framework that integrates retrieval-augmented generation (RAG) with VLMs, introducing four key innovations: a four-dimensional affordance descriptor, a hierarchical affordance-aware RAG module, a scene-graph-based spatial constraint reasoner, and a three-tier self-reflective retry mechanism covering 14 failure modes. This approach unifies functional affordances, spatial relationships, and closed-loop recovery within a VLM-based grasping system for the first time. Evaluated across 12 benchmarks spanning single grasps, interactive tasks, and long-horizon missions, the method achieves an overall success rate of 78.3%, representing an absolute improvement of 53.3 percentage points over pure VLM baselines.

affordance awarenesscluttered environmentsfailure recovery

Visual Affordances: Enabling Robots to Understand Object Functionality

May 08, 2025
TA
T. Apicella
🏛️ Istituto Italiano di Tecnologia | Queen Mary University of London | Idiap Research Institute | Ecole Polytechnique Federale de Lausanne

Current visual affordance prediction research suffers from fragmented task definitions—such as grasp detection and affordance classification—each redefining “affordance” independently, leading to incomparable benchmarks and poor reproducibility, thereby hindering generalization in robotic interaction. To address this, we propose a unified modeling framework: (1) formally define visual affordance as a mapping from object visual features to physical attributes (e.g., mass) and subsequently to task-relevant interactions; (2) introduce the *Affordance Sheet*, a standardized documentation protocol specifying data curation, evaluation metrics, and implementation details; and (3) establish a cross-task benchmark and reproducibility diagnostic toolkit integrating visual perception, physical attribute estimation, and interaction modeling. This work is the first to achieve conceptual unification, evaluation consistency, and implementation transparency in affordance prediction, significantly improving model reliability and generalization in real-world robotic settings.

Addressing reproducibility issues in visual affordance prediction methodsBridging affordance perception with physical properties for robot interactionUnifying formulations for diverse affordance tasks like grasping and segmentation

To address the challenge of functional grasping in dexterous robotic tool manipulation, this paper proposes GAAF-Dex, an end-to-end framework that learns granularity-aware affordance features from human-object interactions to jointly localize functional contact regions (fine-grained) and predict dexterous grasp poses (coarse-grained). We introduce a novel weakly supervised cross-view learning paradigm—leveraging exocentric supervision to guide egocentric affordance estimation—and a force-feedback-driven coarse-to-fine grasp post-processing module. The method integrates multi-granularity affordance representation, functional finger coordinate localization, hand-to-end-effector coordinate transformation, and joint modeling of exocentric and egocentric images. Evaluated on our newly constructed FAH dataset (6K images, 18 tools, 6 task categories), GAAF-Dex significantly outperforms state-of-the-art methods in both functional region localization accuracy and dexterous gesture prediction. The code is publicly available.

Locating functional affordance areas via granularity-aware extractionPredicting grasp gestures from hand-object interaction cuesTeaching robots dexterous tool grasping using affordance features

Latest Papers

What's happening recently
View more

This work addresses the limited generalization of existing robotic planning systems, which rely heavily on visual appearance and neglect task-relevant functional properties, thereby struggling in novel robot-object interaction scenarios. To overcome this, the authors propose A4D, a novel approach that establishes a shared latent space centered on functional attributes (e.g., “movable”). A4D enables efficient reasoning through alignment of visual and functional embeddings and proximity-driven matching to functional prototypes. Furthermore, it incorporates an uncertainty-based few-shot mechanism for discovering new functions in previously unseen contexts. The method achieves 94% accuracy in known-function inference—surpassing prior approaches by over 15 percentage points—and, for novel functions, requires less than 10% of the training data to boost accuracy from 70% to above 90%, while also delivering a hundredfold speedup in inference time.

affordance reasoningfunctional latent spacegeneralization

Existing knowledge representation approaches struggle to capture cooperative affordances in multi-agent social contexts—namely, the possibilities through which agents extend their action capabilities via collaboration. This work introduces, for the first time, an extension of the affordance concept to multi-agent cooperative scenarios by proposing a computable ontological framework that formally defines “cooperative affordances.” Building upon ontology engineering principles, the framework establishes composable and extensible basic representational patterns. By combining these patterns, the system effectively models collaborative interactions ranging from elementary to complex, demonstrating both expressive power and practical utility across diverse social settings.

affordancescooperative affordancesknowledge representation

Current vision-language-action (VLA) models struggle to focus on task-relevant functional interaction regions due to their reliance on holistic object appearance, limiting robustness in unstructured environments. This work proposes an implicit affordance injection mechanism that leverages a zero-shot affordance teacher model to extract language-conditioned affordance representations and aligns them with intermediate visual features of the VLA model. This approach internalizes task-oriented affordance awareness without requiring explicit masks or additional modules. While preserving inference efficiency, the method reshapes visual representations to heighten sensitivity to manipulation-critical regions. Experiments demonstrate that the proposed model significantly outperforms strong baselines in both simulated and real-world environments, achieving higher task success rates and improved training efficiency.

affordance representationrobotic manipulationunstructured environments

Existing large models struggle to generate non-obvious yet physically feasible tool-use strategies in open-world settings due to insufficient grounding in visual and physical constraints. To address this, this work introduces MM-CreativityBench, the first benchmark for systematically evaluating embodied creativity in large models. The proposed approach incorporates a functional attribute alignment mechanism that leverages a functional knowledge base for supervision, multi-view structured scene representations, multi-turn interactive reasoning, and preference learning via Direct Preference Optimization. This framework significantly improves the model’s accuracy in selecting appropriate objects and parts, effectively mitigates hallucination and grounding errors, and consistently enhances performance across creative physical reasoning tasks.

affordance groundingcreative problem-solvinglarge multimodal models