日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ヒューマノイド制御arXiv:2609.28378

ForgetMimic: 強化学習によるヒューマノイド制御のための運動アンラーニング

ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control

シェア:XThreadsFacebookLINEはてブBluesky

強化学習で学習したヒューマノイド制御ポリシーから、指定した運動だけを忘却させ、他の運動性能を維持する初の運動レベルアンラーニング手法を提案。

詳しい要約

1. どんなもの?

- 物理世界のヒューマノイド制御において、学習済みポリシーから特定の動作を忘却させる初の動作レベル・アンラーニング手法。 - 対象は強化学習(RL)でN個の動作を学習したポリシーπ_θ。 - 目標はK個の対象動作の性能を劣化させ、残りのN-K個の動作の有効性を保持すること。 - 安全性・プライバシー(悪意ある動作、毒された動作、GDPRの忘れられる権利)への対応を動機とする。

2. 先行研究と比べてどこがすごい?

- 従来のヒューマノイド制御は人間のデモンストレーションを活用したRLで多様な運動を実現してきたが、学習済みポリシーから特定動作を除去する方法は十分に検討されていなかった。 - ForgetMimicは物理世界のヒューマノイド制御に特化した初の動作レベル・アンラーニング手法である点が新しい。 - 既存のアンラーニング研究は画像分類など他分野が中心で、ロボット制御への適用は未開拓だった。

3. 技術・手法の肝は?

- 核心は、N個の動作で訓練されたポリシーπ_θに対し、K個の対象動作の性能を劣化させつつ、残りのN-K個の動作の性能を維持する枠組み。 - ロボット制御におけるアンラーニング失敗を引き起こす2つの重要な訓練メカニズムを特定し、解決する。 - 具体的なアルゴリズムや損失関数の詳細は要旨からは不明。

4. どうやって有効だと検証した?

- Unitree G1およびH2ヒューマノイドロボットを用い、Dance、Fight、Flipを含む12種類の動作で広範な実験を実施。 - 実験結果は、ForgetMimicが指定動作の記憶を効果的に消去しつつ、他の全動作の正常動作を維持することを示した。

5. 議論はある?

- 安全性とプライバシーの観点から、悪意ある動作、毒された動作、最適でない動作の除去、およびGDPRなどの規制における忘れられる権利への対応の重要性を議論。 - アンラーニング失敗を招く訓練メカニズムの特定と解決について言及。 - 限界や今後の課題についての具体的な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、強化学習(RL)によるヒューマノイド制御、人間のデモンストレーションを用いた模倣学習、機械学習におけるアンラーニング(machine unlearning)が挙げられる。 - 同分野の定番として、Unitree G1/H2を用いたロコモーション研究や、DeepMimicなどの物理ベースの動作模倣手法が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xukun Luan, Zhongxiang Lei, Chen Gong, Shaowei Li, Yuanguo Bi, Jinyan Liu

分類: cs.RO, cs.CR, cs.LG

原文アブストラクト

Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains insufficiently explored. Addressing this issue is motivated by pressing safety and privacy concerns: the removal of malicious, poisoned, or suboptimal motions, as well as copyright-protected motions subject to the right to be forgotten under regulations such as the GDPR, is of critical importance. To this end, we propose {ForgetMimic}, the first motion-level unlearning method designed specifically for physical-world humanoid control. The core idea of ForgetMimic is as follows: given a policy $π_θ$ trained on $N$ motions, our method degrades performance on a target subset of $K$ motions while preserving the effectiveness of the remaining $N-K$ motions. Furthermore, we identify and resolve two key training mechanisms in robot control that lead to unlearning failure. We conduct extensive experiments on the Unitree G1 and H2 humanoid robots across 12 motions, including Dance, Fight, Flip, and others. Experimental results demonstrate that ForgetMimic effectively eliminates memory of designated motions while maintaining the normal operation of all other motions.

関連論文

PR本紙発行元 EmplifAI