🤖 AI Summary
This study addresses the challenges of achieving high-speed, precise control in complex hydraulic excavators and the prohibitive costs of real-hardware interaction. We propose a direct online model-based reinforcement learning framework that requires neither simulation pre-training nor demonstration data. The method integrates probabilistic dynamics ensemble modeling with sampling-based model predictive control, and introduces an accuracy-gated contour reward function to balance operational speed with trajectory precision. Experiments conducted on an 11.5-ton excavator demonstrate that the proposed approach attains tracking performance comparable to conventional methods requiring over 100 minutes of interaction within merely 20 minutes. Furthermore, sub-centimeter path errors are achieved under high-speed operating conditions after 40 minutes, significantly enhancing both sample efficiency and control accuracy.
📝 Abstract
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven excavator simulator, the framework achieves higher sample efficiency than the evaluated model-based reinforcement learning baselines. We validate the framework by learning directly on an 11.5-ton Menzi Muck M445 hydraulic excavator, without demonstrations or simulation pretraining. After 20 minutes of interaction, the controller reaches tracking accuracy comparable to prior learned controllers trained on 100-150 minutes of data. After 40 minutes, it sustains sub-centimeter mean path error at high operating speeds.