Score
Designs and implements motion planners, controllers, and task-level coordinators for two manipulators acting together to perform coordinated actions. This includes generating synchronized bimanual trajectories, timed handoffs, dual-arm grasping and unscrewing maneuvers, and coordinated force/pose control to achieve end-to-end bimanual task execution in constrained or cluttered settings.
This work addresses the challenge of efficiently generating executable plans for dual-arm robotic tasks from human demonstrations, where bimanual coordination strategies are often complex and difficult to model. The authors propose a novel approach that leverages a single RGB video demonstration to synthesize structured, modular behavior tree plans. Their method uniquely integrates Shannon information theory to analyze information flow between hands, scene graph parsing to extract action semantics, and one-shot learning to enable generalization. By unifying these components within a behavior tree framework, the approach produces adaptable execution plans without requiring extensive training data. Evaluated on both a newly curated dataset and existing public benchmarks, the method demonstrates significant performance gains over current state-of-the-art techniques, marking a notable advance in centralized bimanual coordination planning.
This work addresses the lack of efficient and intuitive methods for expressing hierarchical structures involving both relative and absolute constraints in remote assembly tasks. The authors propose a bimanual extended reality (XR) interaction technique that enables users to construct nested constraint groups by grasping objects with each hand, assigning either relative constraints—where robot-optimized poses are computed—or absolute constraints—where user-specified poses are fixed—to each group. This approach introduces, for the first time in bimanual XR, hierarchical modeling of mixed 6DoF constraints, augmented with visual convex hull representations to generate robot-interpretable assembly specifications. By minimizing the need for manual specification of precise object poses, the method significantly enhances the efficiency and flexibility of remote teleoperation in complex assembly scenarios.
To address the challenges of poor generalizability and unnatural transitions in human-object collaborative pick-and-place animation generation within cluttered environments, this paper proposes a hierarchical goal-driven framework. First, a bimanual scheduler generates task-critical keyframes; second, a neural implicit planner models hand trajectories adaptively to diverse object geometries and dynamic obstacle configurations; third, a DeepPhase controller—enhanced by Kalman-filter-augmented frequency-domain smoothing—enables multi-target linear dynamical motion control. Our key innovations include the first integration of bimanual coordination with neural implicit planning, and a novel frequency-domain-optimized DeepPhase dynamic control paradigm. Experiments demonstrate significant improvements in task success rate and motion naturalness across photorealistic scenarios featuring geometric heterogeneity, movable containers, and dense layouts, outperforming existing single-arm and static-planning approaches.
Long-horizon collaborative tasks for dual robotic arms face challenges including complex spatiotemporal dependencies among subtasks, difficulty in dynamic action allocation, and limited expressiveness of linear programming formulations. This paper proposes the first LLM-driven DAG-structured task decomposition framework, which automatically parses high-level instructions into directed acyclic graphs (DAGs) encoding dependency constraints, and integrates environment perception to enable real-time, dynamic action allocation and parallel adaptive execution across both arms. The method breaks away from predefined operational paradigms, supporting end-to-end, interpretable, and generalizable collaborative planning. Evaluated on the Dual-Arm Kitchen benchmark, it achieves a 52.8% efficiency gain over single-arm systems, improves success rate by 48% and reduces LLM query count by 84.1% compared to conventional dual-arm planners, significantly enhancing robustness and scalability in complex scenarios.
This work addresses the challenges of low efficiency, poor trajectory quality, and difficulty in feasible pose search when single-arm robots perform precise interference-fit assembly in confined spaces. We propose the first end-to-end dual-arm collaborative assembly framework that automatically generates high-quality assembly strategies using only part CAD models and the target assembly pose. Our approach integrates multi-robot motion planning, CAD-driven modeling, collaborative trajectory optimization, and physics-based simulation. We theoretically demonstrate for the first time that dual-arm collaboration significantly enhances assembly performance, providing formal guarantees on execution time and trajectory accuracy, and derive theoretical bounds on robot cell dimensions. Experiments show that, compared to single-arm baselines, our method reduces average execution time by over 50%, substantially improves trajectory quality, and accelerates feasible pose discovery, with results validated through both simulation and physical experiments.
This study addresses the lack of a unified and portable real-time low-level motion planning interface for heterogeneous collaborative robotic arms. To bridge this gap, the authors propose a lightweight and flexible real-time end-effector trajectory control interface built upon the WinGs Operating Studio middleware. The approach integrates n-th-order polynomial interpolation with a quadratic programming (QP) solver to generate smooth, continuously differentiable trajectories that enable precise control over position, velocity, and acceleration. For the first time, real-time low-level motion planning across multiple brands of collaborative arms is achieved under a single interface, featuring on-the-fly replanning capability and cross-platform compatibility. Experimental validation through offline drawing, dynamic grasping, and cross-arm teleoperation demonstrates significant improvements in system generality, deployment efficiency, and usability.
This work addresses the lack of explicit coordination mechanisms in existing vision-language-action (VLA) models for tightly coupled dual-arm tasks, which undermines behavioral reliability, interpretability, and stability. To overcome this limitation, the authors introduce the Structured Action Experts (SAE) module—the first such integration within a VLA framework—employing shared and residual latent variables to model task-level coordination intent and per-arm execution adjustments, respectively. Coupled with a Latent-Aware Controller (LAC), the approach enables synchronized regulation of dual-arm behavior. The method supports real-time modulation of synchronization strength, execution asymmetry, motion smoothness, and safety constraints. Experiments demonstrate a 27% improvement in success rate on tightly coordinated tasks, a doubling of out-of-distribution performance in real-world scenarios (from 13% to 27%), and up to a 25% reduction in task completion time.
This work addresses the challenge of open-loop grasping under uncertainty in object shape and pose, where poor contact coordination often leads to slippage or failure. The authors propose a tactile feedback–based model predictive controller that enables coordinated multi-contact interaction and adaptive force modulation during both approach and grasp phases. Key innovations include perception-driven phase segmentation, arm–hand协同 compensation for pose errors, and a balanced adaptive force coordination mechanism. By analytically linking contact forces to joint motions, the method remains compatible with diverse grasp pose generation strategies. Evaluated across 15,000 simulations involving 478 objects and eight physical experiments, the approach significantly improves grasp success rates while effectively suppressing unintended object motion.