ResSafe: 残差強化学習によるヒューマノイドの安全フィルタリング
ResSafe: Learning Safety Filtering with Residual Reinforcement Learning for Humanoids
ヒューマノイド制御において、タスク遂行ポリシーと安全補正ポリシーを分離し、残差強化学習で安全フィルタを学習する手法を提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Gechen Qu, Tong Zhang, Bike Zhang, Yen-Jen Wang, Koushil Sreenath, Claire Tomlin, Jason Jangho Choi
分類: cs.RO
原文アブストラクト
Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, learned policies can still generate unsafe actions that lead to instability or falls. In this work, we propose residual reinforcement learning as an implicit safety-filtering mechanism for safe humanoid control. Instead of relying on a single nominal policy to simultaneously balance performance, safety, and robustness, we decouple performance and safety. The nominal policy focuses solely on task performance, while a residual policy learns safety corrections. This decoupling leads to a better performance--safety Pareto trade-off and avoids the need for careful tuning of multiple competing reward terms within a single policy training. We show that the residual policy can act as an implicit safety filter.
関連論文
- ViBe: 知覚型ヒューマノイド全身制御のための視覚行動適応ヒューマノイド制御
- GLoRI: グローバル・ローカル参照相互作用によるヒューマノイド移動操作のための閉ループ全身追跡ヒューマノイド制御
- 学習された停止可能性値によるヒューマノイドの安全停止ヒューマノイド制御
- ADAPT: 俊敏な拡散行動事前分布による堅牢で操縦可能なオンライン文章駆動ヒューマノイド制御ヒューマノイド制御
- StableMimic: 人間らしいスムーズな復帰を実現するヒューマノイド動作追跡 - 追跡分布を超えた学習による構造化された転倒後行動ヒューマノイド制御
- 文脈を考慮したモーション事前分布によるヒューマノイド制御の学習ヒューマノイド制御