サポート保持蒸留によるマルチエージェント協調
Multi-Agent Coordination via Support-Preserving Distillation
オフラインMARLの生成ポリシー蒸留において、教師のノイズ割り当てを最適輸送で修正し、モード間の誤り伝播を防ぐ手法を提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Sangmin Lee, Youngju Na, Chanmi Lee, Sung-eui Yoon
分類: cs.LG, cs.MA, cs.RO
原文アブストラクト
Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples can be routed toward conflicting coordination modes. The teacher then produces samples between valid modes, and because the distillation loss regresses each local actor onto the conditional mean of the teacher's output given local input, this error is not absorbed but propagated to the student. To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training. We additionally study a shared-randomness variant that uses a shared noise component at execution to expose the residual gap intrinsic to strict-product execution. On controlled diagnostics and offline MARL benchmarks, MoSDOT improves endpoint quality and routing consistency, particularly on datasets exhibiting multimodal joint behavior.
関連論文
- 未来の協力者と協調する:参加タイミングがずれるマルチエージェント強化学習マルチエージェント強化学習
- 部分的観測動的ゲームにおける未知の対戦相手に対するレベルK政策の編成マルチエージェント強化学習
- 置換ロバスト性だけでは不十分:マルチエージェントTransformer方策における行動崩壊マルチエージェント強化学習
- テスト時マルチエージェント協調のための分解価値勾配フローマルチエージェント強化学習
- MA-JEPA: マルチエージェント強化学習のための結合埋め込み世界モデルマルチエージェント強化学習
- MATES: 凍結した単一エージェント方策の観測変換によるマルチエージェント相互作用の学習マルチエージェント強化学習