人間とロボットの協調における人間フィードバック強化学習のスコーピングレビューと実験研究
A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration
産業4.0での人間とロボットの協調に向け、人間フィードバック強化学習(RLHF)の研究をレビューし、VR実験でユーザー主導のフィードバックがシステム主導より安全知覚を捉えやすいことを示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Alexandra Coroiu, Andrea Vogt, Viktor Werbilo, Andreas Poppele, Johann Christensen, Sven Hallerbach
分類: cs.HC, cs.AI, cs.RO
原文アブストラクト
Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots. Practical challenges remain regarding safety during AI development, human feedback quality, and bidirectional human-robot adaptation. We conducted a scoping review of RLHF in HRC systems, mapping methods that address these challenges. Following PRISMA guidelines, we screened 199 records and included 20 peer-reviewed publications (2020-2025) spanning multiple HRC domains. To our knowledge, this is the first review focused on the bidirectional, closed-loop design of RLHF. Our review found multiple feedback modalities enabling data collection in various feedback formats. Collected data can be integrated at different stages of AI training, resulting in a multi-step development process. Pilot experiments are commonly used to evaluate HRC systems based on both human and robot metrics. To empirically test a key gap identified in the review, we conducted a between-subjects VR experiment comparing system- and user-initiated feedback on robot proxemic behaviour for safe navigation. Using Bayesian models, we analysed the relation between the collected feedback and safety metrics: psychological safety (post-experiment questionnaire) and physical safety (inverse time-to-collision). Results show that user-initiated feedback captures perceived safety better than system-initiated feedback, indicating that feedback timing directly affects feedback quality. Our review and experiment findings show that RLHF relies on appropriate feedback methods to ensure AI safety in HRC, and future RLHF research should prioritise realistic HRC experiments evaluating the effects of feedback collection methods on relevant human and robot metrics.
関連論文
- 人間とロボットの協調組立における解釈可能な認知作業負荷評価のための注意と行動手がかりを統合した視覚ベースフレームワーク人間ロボット協調
- すべての一致が裏付けとなるわけではない:人間とロボットの協調における型付き行動受理のための来歴保存型多視点融合人間ロボット協調
- 隠れた目標下でのゼロショット人間ロボット協調のための構造化LLM推論人間ロボット協調
- 人間とロボットのチーム成功のための能動的信頼管理:信頼修復から信頼充足の視点へ人間ロボット協調
- プレハブ建設における計画指向の人間ロボット協調のための認知・身体的負担の知覚ダイナミクスの理解とモデル化人間ロボット協調
- 人間とロボットの協調の探求:困難なタスクにおける対話モダリティの分析人間ロボット協調