置換ロバスト性だけでは不十分:マルチエージェントTransformer方策における行動崩壊
Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies
マルチエージェントTransformer方策のエージェント順序置換に対するロバスト性を調べ、低い置換誤差が全エージェント同一行動による見かけのロバスト性である「行動崩壊」を引き起こすことを示し、行動多様性指標の必要性を提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Amit Thakur, Mukesh Singhal
分類: cs.RO, cs.LG, cs.MA
原文アブストラクト
Transformer policies are attractive for multi-agent robot learning because self-attention can model interactions among agents. However, multi-agent teams are unordered, while transformers typically process agents as ordered token sequences. We study how this mismatch affects cooperative navigation policies under agent-order permutations. Our results show that low permutation error alone can be misleading: policies may appear robust simply because all agents choose the same action. We therefore evaluate policies using both permutation-consistency metrics and action-collapse diagnostics, including action diversity, same-action fraction, and maximum action frequency. A PPO-ID baseline yields non-collapsed behavior but remains order-sensitive, while strong equivariance regularization can still induce homogeneous behavior. A weak equivariance penalty improves the robustness while preserving more diverse actions for teams with \(N=3\) agents, whereas teams with \(N=4\) agents require substantially smaller regularization weights. These findings suggest that multi-agent transformer policies should be evaluated not only by return and permutation robustness, but also by whether they maintain non-collapsed, differentiated multi-agent behavior.
関連論文
- テスト時マルチエージェント協調のための分解価値勾配フローマルチエージェント強化学習
- MA-JEPA: マルチエージェント強化学習のための結合埋め込み世界モデルマルチエージェント強化学習
- MATES: 凍結した単一エージェント方策の観測変換によるマルチエージェント相互作用の学習マルチエージェント強化学習
- 完全ビザンチン耐性マルチエージェント強化学習マルチエージェント強化学習
- 山火事対応における自律UAV探査のためのマルチエージェント強化学習マルチエージェント強化学習
- 予測シールディングによる分散型安全マルチエージェント強化学習マルチエージェント強化学習