A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of safety alignment, feedback quality, and bidirectional adaptation in Reinforcement Learning from Human Feedback (RLHF) within human–machine collaboration. It is the first to investigate the bidirectional closed-loop mechanism of RLHF. Methodologically, a systematic review was conducted following PRISMA guidelines, complemented by empirical analyses integrating virtual reality experiments with Bayesian modeling. The results demonstrate that feedback timing significantly influences interaction quality. Furthermore, compared to system-initiated prompts, user-initiated feedback more precisely captures psychological safety and enhances overall system security. By establishing both a theoretical foundation and empirical evidence, this work provides critical insights for optimizing RLHF frameworks in collaborative human–machine systems.
📝 Abstract
Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots. Practical challenges remain regarding safety during AI development, human feedback quality, and bidirectional human-robot adaptation. We conducted a scoping review of RLHF in HRC systems, mapping methods that address these challenges. Following PRISMA guidelines, we screened 199 records and included 20 peer-reviewed publications (2020-2025) spanning multiple HRC domains. To our knowledge, this is the first review focused on the bidirectional, closed-loop design of RLHF. Our review found multiple feedback modalities enabling data collection in various feedback formats. Collected data can be integrated at different stages of AI training, resulting in a multi-step development process. Pilot experiments are commonly used to evaluate HRC systems based on both human and robot metrics. To empirically test a key gap identified in the review, we conducted a between-subjects VR experiment comparing system- and user-initiated feedback on robot proxemic behaviour for safe navigation. Using Bayesian models, we analysed the relation between the collected feedback and safety metrics: psychological safety (post-experiment questionnaire) and physical safety (inverse time-to-collision). Results show that user-initiated feedback captures perceived safety better than system-initiated feedback, indicating that feedback timing directly affects feedback quality. Our review and experiment findings show that RLHF relies on appropriate feedback methods to ensure AI safety in HRC, and future RLHF research should prioritise realistic HRC experiments evaluating the effects of feedback collection methods on relevant human and robot metrics.
Problem

Research questions and friction points this paper is trying to address.

Human-Robot Collaboration
Reinforcement Learning from Human Feedback
Safety
Feedback Quality
Bidirectional Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning from Human Feedback
Human-Robot Collaboration
Bidirectional Closed-loop Design
Bayesian Models
Proxemic Behaviour
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alexandra Coroiu
DLR Institute for AI Safety and Security, Wilhelm-Runge-Straße 10, 89081, Ulm, Germany
A
Andrea Vogt
DLR Institute for AI Safety and Security, Wilhelm-Runge-Straße 10, 89081, Ulm, Germany
V
Viktor Werbilo
DLR Institute for AI Safety and Security, Rathausallee 12, 53757, Sankt Augustin, Germany
A
Andreas Poppele
DLR Institute for AI Safety and Security, Wilhelm-Runge-Straße 10, 89081, Ulm, Germany
J
Johann Christensen
DLR Institute for AI Safety and Security, Rathausallee 12, 53757, Sankt Augustin, Germany
S
Sven Hallerbach
DLR Institute for AI Safety and Security, Rathausallee 12, 53757, Sankt Augustin, Germany