日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチエージェント強化学習/全身制御arXiv:2610.01102

MASkillBlender: スキルブレンディングによるマルチヒューマノイド全身協調の分散制御

MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済みの単一ヒューマノイドスキルを再利用し、タスクレベルの報酬だけで複数ヒューマノイドの分散的な全身協調動作を学習するフレームワークを提案。

詳しい要約

1. どんなもの?

- 複数の humanoid が協調して loco-manipulation を行うための分散型 whole-body 制御フレームワーク。 - 高次元 whole-body 制御、分散意思決定、スケーラビリティの課題に対処。 - 再利用可能な pre-trained single-humanoid skills 上で共有 high-level policy を学習。 - タスクレベル報酬のみで協調行動を実現し、タスク固有の motion reference を不要とする。

2. 先行研究と比べてどこがすごい?

- 従来の reinforcement learning は single-humanoid whole-body 制御に焦点。 - マルチ humanoid 設定への拡張は報酬設計やタスク固有設計が大変。 - 提案手法はタスク固有の motion reference なしで協調を学習可能。 - 均質マルチ humanoid 系に対する permutation-based data augmentation を導入し、学習効率を改善。

3. 技術・手法の肝は?

- 共有された分散 high-level policy を pre-trained single-humanoid skills 上で学習。 - タスクレベル報酬のみを使用し、協調行動を獲得。 - 均質マルチ humanoid 系のための permutation-based data augmentation を導入。 - 理論的に、homogeneous Markov game 定式化の下で置換サンプルが元のサンプルの policy-gradient 方向を保持することを示す。

4. どうやって有効だと検証した?

- 2 種類の humanoid embodiment にわたる複数のマルチ humanoid 協調タスクで評価。 - シミュレーション結果により、提案フレームワークが一貫して高いタスク性能を達成。 - 異なるタスクと humanoid embodiment にわたって協調行動を実現できることを示す。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、single-humanoid whole-body control の reinforcement learning 研究、multi-agent reinforcement learning、homogeneous Markov game の理論が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yifan Hu, Luhang Hong, Mingkang Long, Danning Wang, Chengfeng Jia, Rong Su, Junjie Fu, Guanghui Wen

分類: cs.RO, cs.LG, cs.MA

原文アブストラクト

Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.

PR本紙発行元 EmplifAI