🤖 AI Summary
This study addresses the limitations of existing robot control strategies, which rely heavily on domain knowledge, lack transparency, and struggle to unify robust control with learning paradigms. To overcome these challenges, this work proposes an end-to-end quadratic programming (QP) policy framework that, for the first time, unifies robust control, multilayer perceptrons, and robot learning within a single QP abstraction. By integrating domain randomization with black-box modeling, the approach enables automatic parameter tuning without requiring explicit system dynamics. The primary contribution lies in reconciling model interpretability with data-driven flexibility. Both simulation and hardware experiments demonstrate the broad applicability of the proposed method, revealing significant improvements in system robustness and efficient policy optimization under unknown dynamics.
📝 Abstract
We present an end-to-end QP-based policy framework that enables systematic policy construction with minimal domain-specific design, while preserving the transparency and interpretability of model-based control. The proposed policy representation supports both domain-randomized model-based auto-tuning, where policy parameters are optimized over distributions of disturbances and model variations, and black-box policy construction, where the problem is formulated in terms of a (possibly) unknown model without requiring explicit notions of states, inputs, or the underlying system dynamics. We establish connections to existing policy representations and control paradigms, including robust control, multilayer perceptrons, and robot learning, and interpret our end-to-end QP policies as a common abstraction of these approaches. We validate the resulting framework in both simulation and hardware, demonstrating a broad range of applications spanning robustness, automatic policy tuning, and control under unknown system dynamics.