安全なシミュレーションから実世界への転移の証明可能な保証
Provably Safe Sim-to-Real Transfer
シミュレータで訓練した方策を実世界に安全に転移するためのアルゴリズムを提案し、シミュレータと実世界の差に応じたサンプル複雑性の理論的保証を与えた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Tingting Ni, Maryam Kamgarpour
分類: cs.LG, cs.AI
原文アブストラクト
To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where samples are cheap, and then deploy the learned policy in the real world with the hope that it generalizes effectively. Such direct sim-to-real transfer is not guaranteed to succeed: simulator-trained policies can be suboptimal in the real world due to sim-to-real mismatch. Correcting this mismatch requires collecting data from the real system, but in many applications, such as robotics and healthcare, this data-collection process is itself subject to safety constraints. This gives rise to the problem of safe sim-to-real transfer: how can an agent exploit an imperfect simulator while ensuring safe real-world data collection and learning a near-optimal feasible policy for the target system? We address this problem by formulating safe sim-to-real transfer within the framework of reward-free safe RL. We design a computationally efficient algorithm that exploits simulator information to provably reduce real-world interaction while ensuring safe exploration and enabling the computation of a near-optimal feasible policy for any potential reward function. Our real-world sample complexity bound characterizes the benefit of using the simulator in terms of the sim-to-real mismatch.
関連論文
- タスク関連特徴ダイナミクスの忠実度がロボット超音波走査のゼロショットsim-to-real転送を可能にするsim2real
- 実2シミュレーション動力学推定と強化学習によるトルク制御ロボットのSim2Real転送の強化sim2real
- シミュレーションから実世界への性能証明書のためのベッティングsim2real
- 触覚シミュレーションを使わない触覚Sim2Real:ボトルネック潜在再構成による実現sim2real
- R2S-EGO: スパースキャプチャ実世界からシミュレーションへのデュアルプロキシ精緻化sim2real
- LyEvO: リアプノフ誘導進化最適化による安全で堅牢なSim-to-Realポリシー学習sim2real