日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチエージェント強化学習arXiv:2609.26010

MATES: 凍結した単一エージェント方策の観測変換によるマルチエージェント相互作用の学習

MATES: Learning Multi-Agent Interactions by Transforming Observations for Frozen Single-Agent Policies

シェア:XThreadsFacebookLINEはてブBluesky

マルチエージェント環境の観測を、凍結した単一エージェント方策が扱える形式に変換する小さなアダプタを学習することで、方策を再学習せずに協調行動を実現する手法。

詳しい要約

1. どんなもの?

- 複数エージェント強化学習(MARL)において、単一エージェント用に学習済みのポリシーを凍結したまま、マルチエージェント環境で協調行動を実現する入力側適応フレームワーク MATES を提案。 - マルチエージェント観測が単一タスク情報を保持しつつ、隣接エージェント情報を別個に識別可能な構造を持つ問題を対象とする。 - 小さなアダプタを学習し、マルチエージェント観測を凍結ポリシーが期待する形式に変換して行動を誘導する。 - ポリシーの内部アーキテクチャや MARL アルゴリズムの目的関数・更新手順は変更しない。

2. 先行研究と比べてどこがすごい?

- 従来の MARL はゼロから分散ポリシーを学習し、個々のタスク能力と協調を同時に獲得する必要があった。 - MATES は凍結した単一エージェントポリシーを再利用し、入力変換のみを学習するため、学習パラメータ数がフルポリシー学習の 3.5-7.3% で済む。 - ゼロからの MARL 学習を一貫して上回り、フルファインチューニングの性能に迫る。 - デモンストレーションを用いたベースラインとも総合的に競争力があり、学習時に遭遇しないチームサイズでも強いタスク性能を維持する。

3. 技術・手法の肝は?

- マルチエージェント観測を入力とし、凍結された単一エージェントポリシーが期待する観測形式にマッピングする小さなアダプタを学習する。 - アダプタはマルチエージェント経験から学習され、共有環境に適した行動を誘導する。 - 凍結ポリシーの内部アーキテクチャは変更せず、基盤となる MARL アルゴリズムの目的関数と更新手順を保持する。 - オン・オフポリシー両アルゴリズムに適用可能。

4. どうやって有効だと検証した?

- lifelong pathfinding、navigation、cooperative discovery のタスクで評価。 - 離散・連続の観測空間と行動空間を含む。 - オン・オフポリシー両アルゴリズムで MATES を評価。 - フルポリシー学習、ゼロからの MARL 学習、フルファインチューニング、デモンストレーションベースラインと比較。 - 学習時に遭遇しないチームサイズでも性能を検証。

5. 議論はある?

- 観測構造が単一タスク情報を保持しつつ隣接情報を別個に識別可能である場合に、個々の能力を符号化するポリシーを変更せずに効果的なマルチエージェント行動を学習できる証拠を示す。 - ただし、この観測構造を満たさない問題への適用可能性や限界については要旨からは不明。 - アダプタの設計や学習安定性、他の MARL アルゴリズムへの一般化については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:MARL のゼロからの学習、フルファインチューニング、デモンストレーションベースのベースライン。 - 関連手法:オン・オフポリシー MARL アルゴリズム(具体的名称は要旨からは不明)。 - 同分野の定番:QMIX、MADDPG、COMA、MAPPO など。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Elie Abboud, Oren Gal

分類: cs.MA, cs.RO, eess.SY

原文アブストラクト

Multi-agent reinforcement learning (MARL) commonly trains decentralized policies from scratch, requiring agents to acquire individual task competence and coordination simultaneously. Yet many multi-agent problems admit a compatible single-agent counterpart in which the underlying task can be learned in isolation. We introduce Multi-Agent Observation Transformation for Existing Single-Agent Policies (MATES), an input-side adaptation framework for tasks whose multi-agent observations preserve the solo-task information while exposing separately identifiable neighbor information. From multi-agent experience, MATES learns a small adapter that maps this observation into the format expected by a frozen single-agent policy, inducing actions suited to the shared environment without updating the single-agent policy itself. MATES leaves the pretrained policy's internal architecture unchanged and retains the objectives and update procedures of the underlying MARL algorithm. We evaluate MATES using both on- and off-policy algorithms on lifelong pathfinding, navigation, and cooperative discovery, spanning discrete and continuous observation and action spaces. Across all evaluated settings, MATES optimizes only 3.5-7.3% as many parameters as full-policy training while consistently outperforming MARL training from scratch. It approaches the performance of full fine-tuning, remains competitive overall with demonstration-based baselines, and retains strong task performance at team sizes not encountered during training. These results provide evidence that, under this observation structure, effective multi-agent behavior can be learned without modifying the policy that encodes individual competence.

関連論文

PR本紙発行元 EmplifAI