🤖 AI Summary
This study addresses the lack of coupled control and insufficient real-time performance of model predictive control (MPC) in hybrid rigid-pneumatic robotic manipulators. We propose a novel coupled state-space modeling approach to quantify inter-joint coupling effects, alongside a distillation framework that transfers MPC capabilities into a neural policy. By leveraging behavioral cloning and the DAgger algorithm, a lightweight policy network is trained to achieve both high performance and real-time execution. Experimental results demonstrate that the proposed method improves trajectory tracking accuracy by 2.5 times, achieves a goal success rate exceeding 93%, and satisfies a strict 5-millisecond real-time control requirement. This work establishes a new paradigm for the efficient control of hybrid-actuated manipulators.
📝 Abstract
Hybrid manipulators combine motorized rigid joints with pressure-actuated origami segments. Published arms of this kind are controlled with decoupled per-DOF loops, and the cost of this approximation has not been quantified, because the coupled model needed to measure it has not been built. This paper derives such a model for a chain of $N$ alternating revolute joints and Kresling origami segments, including pneumatic chamber dynamics and crease hysteresis. Using the model, we measure the coupling directly and show that its strength varies joint by joint, and that decoupled control loses precisely on the strongly coupled joints while remaining competitive on the one nearly decoupled joint. Coupled model-based controllers track $2.5\times$ tighter than a decoupled PID baseline at lower torque. However, the model predictive controller (MPC) is too slow for real time, and model-free reinforcement learning stalls far below acceptable success rates on a strict settling metric. We therefore distill the MPC into a small neural policy with behavior cloning and DAgger. The distilled policy settles 93-94$\%$ of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not. Where the teacher itself fails, we trace the failure to a limit cycle with the bellows' lightly damped mode, and we remove it by selecting goal postures holdable at low pressure.