未来の協力者と協調する:参加タイミングがずれるマルチエージェント強化学習
Cooperating with Future Collaborators: Multi-Agent RL under Staggered Participation
エージェントの参加タイミングがずれる協調マルチエージェント強化学習において、早期エージェントが将来有用な情報を残し後続エージェントがそれを活用するための訓練手法SPLを提案し、複数環境で性能を向上させた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jianglin Qiao, Siyi Hu, Thien Hoang Nguyen, Zehong Cao, Salah Sukkarieh
分類: cs.AI
原文アブストラクト
In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents participating later. We study this setting as staggered participation (SP), which introduces a cross-time, cross-agent learning dependency because an early action may affect the return through the information it provides and the later policy that uses it. Learning under SP therefore requires both identifying what information is useful for future decisions and learning how later agents should use it. We propose Staggered Participation Learning (SPL), a training-time augmentation that addresses these two parts with prospective acquisition supervision for earlier agents and outcome-supervised receiver learning for later agents. We evaluate SPL across multiple policy-based MARL backbones, environments, and staggered-participation patterns. Across 60 MPE/RWARE backbone setting comparisons, SPL achieves higher observed mean task completion in every case, with an average difference of 14.1%. The gains also extend to eight-agent teams and a physics-based UAV-UGV environment in Isaac Lab, providing evidence across algorithmic, temporal, and embodied settings.
関連論文
- 部分的観測動的ゲームにおける未知の対戦相手に対するレベルK政策の編成マルチエージェント強化学習
- 置換ロバスト性だけでは不十分:マルチエージェントTransformer方策における行動崩壊マルチエージェント強化学習
- テスト時マルチエージェント協調のための分解価値勾配フローマルチエージェント強化学習
- MA-JEPA: マルチエージェント強化学習のための結合埋め込み世界モデルマルチエージェント強化学習
- MATES: 凍結した単一エージェント方策の観測変換によるマルチエージェント相互作用の学習マルチエージェント強化学習
- 完全ビザンチン耐性マルチエージェント強化学習マルチエージェント強化学習