テスト時マルチエージェント協調のための分解価値勾配フロー
Test-time Multi-agent Coordination by Decomposed Value Gradient Flow
オフラインMARLにおいて、生成モデルと価値関数をテスト時のアクション洗練で統合し、Stein変分勾配降下で行動サンプルを高価値領域へ輸送するSCOUTを提案。離散・連続ベンチマークで最高性能を達成。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Dongsu Lee, Haoran Xu, Amy Zhang
分類: cs.LG, cs.RO
原文アブストラクト
Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via Stein variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation. Empirically, SCOUT achieves the best average performance across discrete and continuous offline MARL benchmarks and yields performance improvements in all offline-to-online configurations.
関連論文
- 置換ロバスト性だけでは不十分:マルチエージェントTransformer方策における行動崩壊マルチエージェント強化学習
- MA-JEPA: マルチエージェント強化学習のための結合埋め込み世界モデルマルチエージェント強化学習
- MATES: 凍結した単一エージェント方策の観測変換によるマルチエージェント相互作用の学習マルチエージェント強化学習
- 完全ビザンチン耐性マルチエージェント強化学習マルチエージェント強化学習
- 山火事対応における自律UAV探査のためのマルチエージェント強化学習マルチエージェント強化学習
- 予測シールディングによる分散型安全マルチエージェント強化学習マルチエージェント強化学習