Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

📅 2026-02-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates how vision foundation models can achieve genuine understanding of object affordances by jointly modeling geometric structure and interactive behavior. It identifies, for the first time, geometric perception and interaction perception as two composable fundamental components of affordance understanding. To this end, the authors propose a novel zero-shot fusion strategy that requires no additional training: part-level geometric prototypes are extracted using DINO, and then fused with verb-conditioned spatial attention maps generated by Flux. Experimental results demonstrate that this approach achieves performance comparable to weakly supervised methods under zero-shot settings, thereby validating the effectiveness and novelty of the proposed mechanism.

Technology Category

Computer Vision: Diffusion Models for VisionKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Web query analysis, representation and understanding
📝 Abstract
What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actions engage with those parts. To test this hypothesis, we conduct a systematic probing of Visual Foundation Models (VFMs). We find that models like DINO inherently encode part-level geometric structures, while generative models like Flux contain rich, verb-conditioned spatial attention maps that serve as implicit interaction priors. Crucially, we demonstrate that these two dimensions are not merely correlated but are composable elements of affordance. By simply fusing DINO's geometric prototypes with Flux's interaction maps in a training-free and zero-shot manner, we achieve affordance estimation competitive with weakly-supervised methods. This final fusion experiment confirms that geometric and interaction perception are the fundamental building blocks of affordance understanding in VFMs, providing a mechanistic account of how perception grounds action.
Problem

Research questions and friction points this paper is trying to address.

affordance
geometric perception
interaction perception
Visual Foundation Models
part-level structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

affordance reasoning
geometric perception
interaction perception
visual foundation models
zero-shot fusion
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Q
Qing Zhang
The Australian National University
X
Xuesong Li
The Australian National University, CSIRO
J
Jing Zhang
The Australian National University