Deployable Vision-driven UAV River Navigation via Human-in-the-loop Preference Alignment

📅 2025-11-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Vision-based autonomous drones face deployment challenges in river navigation due to distributional shift and safety risks under real-world dynamics. Method: We propose SPAR-H, a human-in-the-loop (HITL) state-level preference alignment framework that integrates direct preference optimization (DPO) with a dual-path reward-based learning paradigm. It combines an online reward estimator, trust-region policy updates, and an imitation-reinforcement hybrid training strategy to enable efficient online adaptation. Contribution/Results: SPAR-H achieves the highest episode reward and lowest reward variance using only five human-in-the-loop rollouts. Its reward model generalizes robustly to unseen states, substantially mitigating distributional shift. Extensive experiments in real river environments demonstrate SPAR-H’s feasibility for continuous state-level preference alignment and its effectiveness in enhancing navigation safety and adaptability. The framework establishes a novel deployable paradigm for vision-driven adaptive learning in autonomous aerial systems.

Technology Category

Humans and AI: Learning Human Values and PreferencesIntelligent Robots: Learning & Optimization for ROBSearch and Optimization: Learning to Search

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Rivers are critical corridors for environmental monitoring and disaster response, where Unmanned Aerial Vehicles (UAVs) guided by vision-driven policies can provide fast, low-cost coverage. However, deployment exposes simulation-trained policies with distribution shift and safety risks and requires efficient adaptation from limited human interventions. We study human-in-the-loop (HITL) learning with a conservative overseer who vetoes unsafe or inefficient actions and provides statewise preferences by comparing the agent's proposal with a corrective override. We introduce Statewise Hybrid Preference Alignment for Robotics (SPAR-H), which fuses direct preference optimization on policy logits with a reward-based pathway that trains an immediate-reward estimator from the same preferences and updates the policy using a trust-region surrogate. With five HITL rollouts collected from a fixed novice policy, SPAR-H achieves the highest final episodic reward and the lowest variance across initial conditions among tested methods. The learned reward model aligns with human-preferred actions and elevates nearby non-intervened choices, supporting stable propagation of improvements. We benchmark SPAR-H against imitation learning (IL), direct preference variants, and evaluative reinforcement learning (RL) in the HITL setting, and demonstrate real-world feasibility of continual preference alignment for UAV river following. Overall, dual statewise preferences empirically provide a practical route to data-efficient online adaptation in riverine navigation.
Problem

Research questions and friction points this paper is trying to address.

Addressing simulation-to-reality gaps in vision-based UAV river navigation
Aligning autonomous policies with human safety and efficiency preferences
Enabling data-efficient adaptation through limited human interventions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-in-the-loop preference alignment for UAV navigation
Fuses direct preference optimization with reward-based training
Enables data-efficient online adaptation via statewise preferences
🔎 Similar Papers
2024-07-11IEEE/RJS International Conference on Intelligent RObots and SystemsCitations: 2
💼 Related Jobs
No related jobs found.
Z
Zihan Wang
School of Mechanical Engineering, Purdue University, West Lafayette, IN 47907, USA
J
Jianwen Li
School of Mechanical Engineering, Purdue University, West Lafayette, IN 47907, USA
Li-Fan Wu
Li-Fan Wu
School of Mechanical Engineering, Purdue University, West Lafayette, IN 47907, USA
Nina Mahmoudian
Nina Mahmoudian
Professor of Mechanical Engineering, Purdue University
Field RoboticsMarine RoboticsMobile RoboticsMulti-Robot SystemsCyber-Physical Systems