🤖 AI Summary
To address the sim-to-real transfer challenge in high-reduction-ratio robotic systems, this paper proposes a physics-guided gain regularization and parameter-conditioned joint training framework. Methodologically, it models PID gains as proxies for unmodeled dynamics; leverages empirically calibrated proportional gains from real-world experiments to constrain the local input–output sensitivity of an RNN-based neural controller; and enhances policy generalization via system-parameter conditioning. The approach integrates domain randomization, sensitivity analysis, and physics-informed regularization loss. Evaluated on a commercial two-wheeled balancing robot, it achieves highly consistent angular motion responses between simulation and reality—improving convergence time alignment and significantly suppressing oscillations. Crucially, the method requires neither high-fidelity dynamical modeling nor ideal (e.g., backdrivable) actuators, making it particularly suitable for low-cost, non-backdrivable hardware platforms. It quantitatively narrows the sim-to-real gap while maintaining practical deployability.
📝 Abstract
Simulation-to-real transfer using domain randomization for robot control often relies on low-gear-ratio, backdrivable actuators, but these approaches break down when the sim-to-real gap widens. Inspired by the traditional PID controller, we reinterpret its gains as surrogates for complex, unmodeled plant dynamics. We then introduce a physics-guided gain regularization scheme that measures a robot's effective proportional gains via simple real-world experiments. Then, we penalize any deviation of a neural controller's local input-output sensitivities from these values during training. To avoid the overly conservative bias of naive domain randomization, we also condition the controller on the current plant parameters. On an off-the-shelf two-wheeled balancing robot with a 110:1 gearbox, our gain-regularized, parameter-conditioned RNN achieves angular settling times in hardware that closely match simulation. At the same time, a purely domain-randomized policy exhibits persistent oscillations and a substantial sim-to-real gap. These results demonstrate a lightweight, reproducible framework for closing sim-to-real gaps on affordable robotic hardware.