日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
群制御arXiv:2610.06290

GAMBIT: 連続マルチロボット軌道計画の学習

GAMBIT: Learning to Plan Continuous Multi-Robot Trajectories

シェア:XThreadsFacebookLINEはてブBluesky

マルチロボットの軌道実行において、模倣学習と強化学習を組み合わせて協調的な動作プリミティブ選択を学習し、衝突回避を保証するフレームワークを提案。

詳しい要約

1. どんなもの?

- 複数ロボットの連続軌道計画を学習するフレームワークGAMBITを提案。 - 個々のロボットが局所的な報酬最大化を犠牲にしてチーム全体の性能を向上させる「自己犠牲的行動」を学習。 - double-integrator連続ダイナミクスに焦点を当て、motion primitiveの協調選択を模倣学習と強化学習で獲得。 - 衝突回避を保証するsafeguarded rollout mechanismを導入。 - 集中型プランナや分散型リアクティブプランナを上回り、1000台以上のロボットを数百ミリ秒未満の遅延で調整可能。

2. 先行研究と比べてどこがすごい?

- 手設計のヒューリスティクスでは捉えにくい自己犠牲的行動を学習可能。 - 集中型motion plannerや分散型reactive plannerを含む複数のベースラインを大幅に上回る性能。 - 強力なスケーラビリティを示し、1000台以上のロボットを連続領域で調整。 - 計画遅延が数百ミリ秒未満と高速。

3. 技術・手法の肝は?

- 協調的なmotion primitive選択を模倣学習で初期学習し、その後強化学習でファインチューニング。 - safeguarded rollout mechanismとbackup trajectoriesにより、常に衝突のない実行を保証。 - double-integrator連続ダイナミクスを対象。

4. どうやって有効だと検証した?

- 実験により、GAMBITが集中型motion plannerや分散型reactive plannerを含む複数のベースラインを大幅に上回ることを示した。 - 1000台以上のロボットを数百ミリ秒未満の計画遅延で調整できるスケーラビリティを実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:centralised motion planners, decentralised reactive planners。 - 関連手法:imitation learning, reinforcement learning, motion primitives, safeguarded rollout mechanism。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Rishabh Jain, Akmaral Moldagalieva, Lorenzo Magnino, Michael Amir, Keisuke Okumura, Ajay Shankar, Wolfgang Hönig, Amanda Prorok

分類: cs.RO, cs.AI, cs.LG, cs.MA

原文アブストラクト

GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to improve overall team performance. Such self-sacrificial behaviours are difficult to capture with manually designed heuristics, particularly in dense, interaction-rich environments. Focusing on double-integrator continuous dynamics, this work studies how to learn such coordinated heuristics over motion primitives for multi-robot trajectory execution. Our framework, GAMBIT, first learns coordinated motion-primitive selection through imitation learning and subsequently fine-tunes the policy through reinforcement learning. We further introduce a safeguarded rollout mechanism with backup trajectories that guarantees collision-free execution at all times. Experiments demonstrate that GAMBIT substantially outperforms a range of baselines, including centralised motion planners and decentralised reactive planners, while exhibiting strong scalability. In particular, it coordinates over a thousand robots with planning latency below a few hundred milliseconds in continuous domains.

関連論文

PR本紙発行元 EmplifAI