日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチエージェント強化学習arXiv:2610.02554

テスト時マルチエージェント協調のための分解価値勾配フロー

Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

シェア:XThreadsFacebookLINEはてブBluesky

オフラインMARLにおいて、生成モデルと価値関数をテスト時のアクション洗練で統合し、Stein変分勾配降下で行動サンプルを高価値領域へ輸送するSCOUTを提案。離散・連続ベンチマークで最高性能を達成。

詳しい要約

1. どんなもの?

- オフラインMARLにおける表現力と価値最適化のトレードオフを解決するSCOUTを提案。 - 生成基盤モデルと学習済み価値関数をテスト時アクションリファインメントで統合。 - 初のオフラインMARLフレームワークで、flow-matching behavioral priorとdecomposed value functionを分離訓練。 - テスト時にStein variational gradient descentで行動サンプルを高価値領域へ輸送。 - 輸送ステップ数が適応的テスト時スケーリングを制御し、固定正則化係数を置換。

2. 先行研究と比べてどこがすごい?

- 従来の生成的ポリシーは多峰性調整を表現できるが高価値領域を識別できない。 - 価値最適化ポリシーはQ関数を活用するが多峰性を単一支配モードに崩壊させる。 - SCOUTは両者をテスト時リファインメントで統合し、モード崩壊を回避しつつ高価値領域を探索。 - 単一エージェントのモード崩壊が共同調整を破壊する問題に対処。 - 同時ドリフトによる未見行動空間への逸脱を防ぐ。

3. 技術・手法の肝は?

- 二つの分離コンポーネント:flow-matching behavioral priorとdecomposed value functionを訓練。 - テスト時にStein variational gradient descentで行動サンプルを高価値領域へ輸送。 - 輸送ステップ数が適応的テスト時スケーリングを制御。 - IGM原理の下で、輸送収束に伴い消滅する単一項KL boundを証明。 - 既約加算残差はIGM違反に比例。

4. どうやって有効だと検証した?

- 離散および連続オフラインMARLベンチマークで最高平均性能を達成。 - すべてのオフライン-to-オンライン構成で性能改善を示す。 - 理論的保証:IGM下でのKL boundと残差の比例性を証明。

5. 議論はある?

- 理論的保証はIGM原理に依存し、IGM違反が残差に影響。 - 輸送ステップ数による適応的スケーリングの有効性を実証。 - モード崩壊と未見領域へのドリフトを同時に解決する枠組みを提供。 - 要旨からは計算コストやスケーラビリティの詳細は不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:オフラインMARLにおける生成的ポリシー、価値最適化ポリシー、flow-matching、Stein variational gradient descent、IGM原理。 - 関連手法:Q-learning、behavior cloning、generative adversarial imitation learning。 - 同分野の定番:MADDPG、QMIX、BCQ、CQL。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dongsu Lee, Haoran Xu, Amy Zhang

分類: cs.LG, cs.RO

原文アブストラクト

Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via Stein variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation. Empirically, SCOUT achieves the best average performance across discrete and continuous offline MARL benchmarks and yields performance improvements in all offline-to-online configurations.

関連論文

PR本紙発行元 EmplifAI