A Unified Dynamics Framework for Reinforcement Learning and Classical Control of a Six-DOF Pipeline-Tracking ROV in NVIDIA Isaac Sim

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of reinforcement learning controllers for underwater robots caused by discrepancies between training and deployment physics models. We propose, for the first time, a dual-branch unified dynamics architecture based on a single USD scene. This framework enables NumPy simulations and Isaac Sim to share identical ocean equation parameters, ensuring physical consistency and facilitating fair comparisons between reinforcement learning and classical control algorithms in high-fidelity environments. Experimental results demonstrate that under Coriolis-force-inclusive dynamics, PPO, TRPO, and feedback linearization all achieve 100% success rates with optimal tracking accuracy. Furthermore, ablation studies quantify the impact of Coriolis forces on stability, revealing an approximately 19% reduction in RMS error for TRPO and validating the competitiveness of classical control methods in specific scenarios.
📝 Abstract
Reinforcement learning controllers for underwater vehicles are usually trained against one physics representation and deployed against another, so reported performance does not always describe behavior outside training. This paper presents a pipeline-tracking architecture for a six-degree-of-freedom remotely operated vehicle (ROV) in which one Universal Scene Description (USD) scene supplies the real BlueROV2-Heavy mass, added-mass, damping, buoyancy, and thruster parameters to both halves of the system: a vectorized NumPy implementation of Fossen's marine-craft equations, and an interactive NVIDIA Isaac Sim deployment applying the identical equations as PhysX forces at every step. The Coriolis-centripetal term is the primary dynamics model in both branches; a controlled ablation on PPO and TRPO shows that including it does not destabilize either algorithm and modestly improves tracking, about 19 percent lower standoff RMS error for TRPO. Five reinforcement learning algorithms, PPO, soft actor-critic, TD3, DDPG, and TRPO, are trained against one environment, reward, and randomized evaluation harness through a checkpoint-compatibility layer scoring any policy with the same code. The pipeline is extended with six classical baselines, PID, sliding-mode, fuzzy logic, feedback linearization, model predictive control, and an adaptive neuro-fuzzy inference system, driven by the same guidance geometry and thruster allocation as the learned policies. Under Coriolis-enabled dynamics, PPO, TRPO, and feedback linearization reach the strongest combination of 100 percent success and competitive accuracy; PID, fuzzy control, and the neuro-fuzzy baseline also reach 100 percent success with looser tracking; DDPG and TD3 each show a specific, explainable failure mode rather than a general weakness of off-policy learning; and classical control remains a strong baseline against the best learned policies.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Underwater ROV
Pipeline Tracking
Dynamics Simulation
Classical Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Dynamics Framework
Reinforcement Learning
Pipeline-Tracking ROV
NVIDIA Isaac Sim
Classical Control Baselines
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Cheng Siong Chin
Faculty of Science, Agriculture, and Engineering, Newcastle University Singapore, Singapore 828608, Singapore
M
M. Venkateshkumar
Amrita School of Engineering, Coimbatore, Amrita Vishwa Vidyapeetham, India
Jianhua Zhang
Jianhua Zhang
Beijing University of Posts and Telecommunications, CHINA
Signal ProcessingWireless CommunicationRadio channel Measurement and ModellingChannel SimulationTerminal Testing