部分的観測動的ゲームにおける未知の対戦相手に対するレベルK政策の編成
Orchestrating Level-$K$ Policies Against Unknown Opponents in Partially-Observable Dynamic Games
未知の対戦相手の推論レベルを推定して応答を選ぶ従来手法を、部分的観測マルコフゲームにおける動的な政策編成問題として定式化し、強化学習ベースの編成器が分類ベースより高いリターンを達成することを示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Addison Kalanther, Sanika Bharvirkar, Daniel Bostwick, Chinmay Maheshwari, Shankar Sastry
分類: cs.MA, cs.RO
原文アブストラクト
Level-$K$ reasoning generates a hierarchy of policies specialized to opponents with different reasoning levels. When an opponent's level is unknown, a common deployment rule estimates that level and selects the corresponding response. In a dynamic, partially-observed game, this selection is repeated, with each choice shaping subsequent states and observations. The response associated with the most likely opponent level need not maximize expected return from the current history. We formulate this deployment problem as dynamic orchestration of a fixed, pretrained policy library in a partially observable Markov game. We compare classification-based orchestrators (CBOs) trained using offline data or on-policy data aggregation with a reinforcement-learning-based orchestrator (RLBO) trained to maximize expected discounted return. In pursuit-evasion experiments, on-policy training improves classification and return, yet RLBO achieves higher return than the on-policy and offline CBOs. Given a pretrained library, RLBO also reaches performance comparable to a policy trained directly over the pursuers' action space with fewer training timesteps. These findings support treating a hierarchy of level-$K$ policies as a resource for orchestration, not a prescription for deployment.
関連論文
- 置換ロバスト性だけでは不十分:マルチエージェントTransformer方策における行動崩壊マルチエージェント強化学習
- テスト時マルチエージェント協調のための分解価値勾配フローマルチエージェント強化学習
- MA-JEPA: マルチエージェント強化学習のための結合埋め込み世界モデルマルチエージェント強化学習
- MATES: 凍結した単一エージェント方策の観測変換によるマルチエージェント相互作用の学習マルチエージェント強化学習
- 完全ビザンチン耐性マルチエージェント強化学習マルチエージェント強化学習
- 山火事対応における自律UAV探査のためのマルチエージェント強化学習マルチエージェント強化学習