Scene-agnostic Hierarchical Bimanual Task Planning via Visual Affordance Reasoning

📅 2025-12-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing dual-arm robot task planning methods for open-world scenarios suffer from a lack of scene priors and struggle to jointly ensure spatial plausibility and effective bimanual coordination. Method: We propose the first scene-agnostic hierarchical planning framework, integrating three novel components: (i) Vision-based Point Grounding (VPG) for grounding actions in 3D space; (ii) Spatial Adjacency–driven Subgoal Planning (BSP) for generating geometrically coherent intermediate goals; and (iii) Interaction Point–guided Bimanual Primitive Selection (IPBP) for coordinated skill invocation. The framework jointly models 3D spatial relations and leverages a structured skill library. Contribution/Results: It requires no scene-specific pretraining or prior knowledge, achieving zero-shot generalization in unseen cluttered environments. It generates compact, semantically meaningful, physically feasible, and highly parallel action sequences. Experiments demonstrate substantial improvements in bimanual coordination and task success rates, establishing a scalable new paradigm for embodied manipulation in open environments.

Technology Category

Intelligent Robots: ManipulationHumans and AI: Human-Aware Planning and Behavior PredictionMultiagent Systems: Multiagent Planning

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic search
📝 Abstract
Embodied agents operating in open environments must translate high-level instructions into grounded, executable behaviors, often requiring coordinated use of both hands. While recent foundation models offer strong semantic reasoning, existing robotic task planners remain predominantly unimanual and fail to address the spatial, geometric, and coordination challenges inherent to bimanual manipulation in scene-agnostic settings. We present a unified framework for scene-agnostic bimanual task planning that bridges high-level reasoning with 3D-grounded two-handed execution. Our approach integrates three key modules. Visual Point Grounding (VPG) analyzes a single scene image to detect relevant objects and generate world-aligned interaction points. Bimanual Subgoal Planner (BSP) reasons over spatial adjacency and cross-object accessibility to produce compact, motion-neutralized subgoals that exploit opportunities for coordinated two-handed actions. Interaction-Point-Driven Bimanual Prompting (IPBP) binds these subgoals to a structured skill library, instantiating synchronized unimanual or bimanual action sequences that satisfy hand-state and affordance constraints. Together, these modules enable agents to plan semantically meaningful, physically feasible, and parallelizable two-handed behaviors in cluttered, previously unseen scenes. Experiments show that it produces coherent, feasible, and compact two-handed plans, and generalizes to cluttered scenes without retraining, demonstrating robust scene-agnostic affordance reasoning for bimanual tasks.
Problem

Research questions and friction points this paper is trying to address.

Planning bimanual manipulation in unseen scenes
Bridging semantic reasoning with 3D-grounded execution
Generating coordinated two-handed actions from visual affordances
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Point Grounding detects objects and interaction points from images.
Bimanual Subgoal Planner creates compact subgoals for coordinated two-handed actions.
Interaction-Point-Driven Bimanual Prompting binds subgoals to synchronized action sequences.
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kwang Bin Lee
Graduate School of Culture Technology, Korea Advanced Institute of Science and Technology (KAIST), 291 Daehak-ro, Yuseong-gu, Daejeon 34141, Republic of Korea
J
Jiho Kang
Graduate School of Culture Technology, Korea Advanced Institute of Science and Technology (KAIST), 291 Daehak-ro, Yuseong-gu, Daejeon 34141, Republic of Korea
Sung-Hee Lee
Sung-Hee Lee
Korea Advanced Institute of Science and Technology (KAIST)
Computer GraphicsCharacter AnimationHuman ModelingTelepresenceRobotics