日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.12140

デジタルツインと人間参加型デモンストレーションを活用した強化学習によるロボットの柔軟性向上

Leveraging Human-In-The-Loop Demonstrations in Reinforcement Learning for Digital Twin-Driven Robot Flexibility

シェア:XThreadsFacebookLINEはてブBluesky

デジタルツインと強化学習、人間のデモンストレーションを組み合わせ、実世界のフィードバックでオンライン適応する枠組みを提案。模倣損失を加えずにデモを活用するデュアルアクター構成で、障害物回避タスクの成功率を大幅に改善した。

詳しい要約

1. どんなもの?

本論文は、collaborative robot の適応性向上を目的に、digital twin (DT)、reinforcement learning (RL)、human demonstrations を組み合わせた human-in-the-loop online training framework を提案する。DT は主にタスク実行前に synthetic data を生成する用途と異なり、camera feeds を通じて物理システムとリアルタイムに同期し、virtual robot が実世界の feedback から observation と policy を更新できる。dual actor framework は imitation learning (IL) を統合するが、RL actor に直接的な imitation loss を加えない。Ufactory Xarm5 collaborative robot で end-effector が障害物を避けつつ目標位置に到達するタスクに適用した。

2. 先行研究と比べてどこがすごい?

従来の DT は主にタスク実行前に synthetic data を生成するために使われるが、本手法の DT は camera feeds を介して物理システムとリアルタイムに同期し、実世界の feedback から virtual robot が observation と policy を更新する点が異なる。また、demonstrations を手動 reprogramming の代わりに適応のガイドとして使えるよう、RL actor に直接的な imitation loss を加えずに IL を統合する dual actor framework を採用している。

3. 技術・手法の肝は?

提案手法の肝は、DT をリアルタイムに物理システムと同期させ、camera feeds から virtual robot の observation と policy を更新する human-in-the-loop online training framework にある。さらに dual actor framework により、RL actor に直接的な imitation loss を加えることなく IL を統合し、demonstrations を適応のガイドとして利用する。これにより、手動 reprogramming を必要とせずに適応できる。

4. どうやって有効だと検証した?

Ufactory Xarm5 collaborative robot を用い、end-effector が障害物を避けながら目標位置に到達するタスクで検証した。物理ワークスペースの変化後に training を再開できることを示した。また、固定された非最適な demonstrations を用いた場合、dual actor framework は actor に imitation loss を加える2手法よりも最終 success rate が大幅に高いことを示した。VR で収集した実 human demonstrations でも同様の傾向が確認され、目標に到達しない demonstrations でも dual actor framework は 83-100% の mean deterministic evaluation success を達成し、imitation-loss 手法の 0-17% を上回った。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究や関連手法として、DT を主にタスク実行前の synthetic data 生成に用いる手法、RL actor に imitation loss を加える2手法、imitation learning (IL)、reinforcement learning (RL)、digital twin (DT) が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuzhu Sun, Mien Van, Nguyen Minh Nhat, Stephen McIlvanna, Sean McLoone

分類: cs.RO, eess.SY

原文アブストラクト

Growing automation makes collaborative robots work in more variable environments, increasing the need for adaptation. We propose a human-in-the-loop online training framework combining a digital twin (DT), reinforcement learning (RL), and human demonstrations. Unlike DTs used mainly to generate synthetic data before task execution, our DT is synchronized with the physical system in real time through camera feeds, allowing the virtual robot to update its observations and policy from real-world feedback. A dual actor framework integrates imitation learning (IL) without adding a direct imitation loss to the RL actor, so demonstrations can guide adaptation instead of manual reprogramming. The proposed framework is demonstrated on the Ufactory Xarm5 collaborative robot, where the robot's end-effector aims to reach the target position while avoiding obstacles. The experiments show that the framework can resume training after a change in the physical workspace and that, with a fixed set of non-optimal demonstrations, the dual actor framework achieves a much higher final success rate than two methods that add an imitation loss to the actor. The same pattern holds with real human demonstrations collected in virtual reality (VR): with demonstrations that never reach the goal, the dual actor framework reached 83-100% mean deterministic evaluation success, against 0-17% for the two imitation-loss methods.

関連論文

PR本紙発行元 EmplifAI