Barrier-Shaped Recurrent Reinforcement Learning for Autonomous Landing on a Heaving Ship Deck

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of precise control and safety assurance for autonomous UAV landing on heaving ship decks by proposing a reinforcement learning method that integrates recurrent neural networks with an asymmetric Actor-Critic architecture. Control Barrier Functions (CBFs) are incorporated into both the reward mechanism and as a runtime safety filter, while large-scale parallel simulation training and visual feedback loops facilitate safe policy learning. Experimental evaluations demonstrate successful landings across 31 motion-capture and 20 vision-based trials, achieving a 77% first-attempt success rate and an average relative contact velocity below 0.3 m/s. These results confirm the method’s effectiveness in ensuring landing safety within complex dynamic environments.
πŸ“ Abstract
This paper addresses autonomous landing of unmanned aerial vehicle (UAV) rotorcrafts on a heaving ship deck using a recurrent policy trained via asymmetric actor-critic reinforcement learning: the critic sees 6 s of future deck height during training, while the actor sees only what the onboard sensors provide in flight, a 17-dimensional state made of the vehicle's position, velocity and attitude relative to the deck and outputs world-frame velocity and yaw-rate commands. Training is performed across 4096 parallel simulated environments, followed by fine-tuning in 16 environments with an onboard vision pipeline in the loop. A control-barrier-function (CBF) stopping margin on the deck-relative vertical state is incorporated at two stages: as a reward term during training, where it halves the median simulated contact speed relative to a policy trained without it, and as a runtime safety filter at deployment, evaluated at every control step to abort and retry the descent when the margin is violated. Because the autopilot's disarm logic cannot detect the vehicle resting on a moving deck, proximity-based thrust cutoff at touchdown is commanded directly in the landing pipeline. The proposed approach is validated using a parallel-manipulator-platform-based deck emulator that reproduces the heaving motion of the ship deck (scaled to 0.70 m peak-to-peak, 7.5 s mean period) and a quadcopter UAV with an onboard camera. Across 31 motion-capture and 20 vision-based trials, the UAV landed every time, with median times to contact of 6.5 and 7.3 s; 77% and 50% landed on the first attempt, with mean deck-relative speeds of 0.29 and 0.27 m/s, respectively, at the instant of thrust cutoff, which is the last speed under the policy's control. Supplementary video: https://youtu.be/S5hDkrSZJt4
Problem

Research questions and friction points this paper is trying to address.

autonomous landing
unmanned aerial vehicle
heaving ship deck
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recurrent Reinforcement Learning
Asymmetric Actor-Critic
Control Barrier Function
Autonomous Landing
Sim-to-Real Transfer
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
R
Ritwik Shankar
Department of Aerospace Engineering, Indian Institute of Technology Kanpur, Kanpur, India
C
Chiranjeev Prachand
Department of Electrical Engineering, Indian Institute of Technology Kanpur, Kanpur, India
Abhishek
Abhishek
Professor, IIT Kanpur
Helicopter dynamicsAeroelasticityUAVs and MAVs
S
Soumya Ranjan Sahoo
Department of Electrical Engineering, Indian Institute of Technology Kanpur, Kanpur, India