🤖 AI Summary
This study addresses the challenges of early training failure and gripper closure timing alignment in whole-body cooperative aerial grasping under partial observability. We propose a recurrent teacher-student framework wherein a teacher policy is trained via privileged reinforcement learning with key-state curriculum learning, then distilled into a vision-based student network relying solely on dual-view point clouds and proprioception. This enables phase-free unified control of flight, arm manipulation, and gripper closure. Additionally, a readiness-sequence-based closure supervision mechanism is introduced to optimize grasp timing. Simulations demonstrate that the proposed policy achieves success rates of 99.97%, 97.14%, and 95.84% under nominal, physics-randomized, and camera-randomized conditions, respectively, with a temporal-spatial alignment error of only 8.12 mm.
📝 Abstract
Aerial grasp-and-lift tasks require whole-body coordination across approach, acquisition, and lifting under partial target observations. Early approach failures can limit exposure to later task stages during training, while changing visibility complicates alignment and closure timing during execution. We present a recurrent teacher-student framework that learns a single policy in simulation to jointly command flight, arm motion, and gripper closure without an explicit task-phase input. A privileged teacher learns through reinforcement learning with a critical-state curriculum that exposes acquisition and lifting states before connecting them to normal approach trajectories. Its behavior is distilled into a recurrent visual student that replaces privileged target states with dual-view point clouds and proprioception, integrating observation history for closed-loop control. A dedicated closure objective supervises closure timing from sustained model-defined readiness sequences. Training and primary evaluation use a simulated acquisition-and-payload model with condition-triggered latching, virtual attachment, and wrench-based payload loading for short-distance lifting. Across 8,996 completed simulation episodes under this model, the frozen student achieves full-task success rates of 99.97%, 97.14%, and 95.84% under nominal, physics/control-randomized, and additional camera-randomized conditions, respectively. The nominal latch-count-weighted mean of per-seed 90th-percentile alignment errors at acquisition is 8.12 mm.