UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of semantic transfer arising from the fragmentation of 2D and 3D affordance perception tasks by proposing a unified framework based on multimodal large language models (MLLMs). Methodologically, the MLLM serves as a shared semantic hub integrated with the SAM pixel decoder and the SONATA point cloud decoder. An innovative token routing mechanism is designed to directly supervise task-specific branches without requiring a language head, thereby facilitating dense prediction optimization across heterogeneous data modalities. Additionally, the UniAfford dataset is constructed to support this research. Experimental results demonstrate that the proposed method achieves state-of-the-art performance under both zero-shot generalization and single-modality protocols, confirming that the routing states effectively convey cross-modal semantics.
📝 Abstract
Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: https://4dvlab.github.io/UniAfford
Problem

Research questions and friction points this paper is trying to address.

affordance perception
2D-3D grounding
multitask learning
embodied interaction
zero-shot generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-Routed Multitask Learning
Affordance Perception
Multimodal Large Language Model
2D-3D Unification
Zero-shot Generalization
🔎 Similar Papers
No similar papers found.