日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチエージェント運転arXiv:2609.35916

VehicleArena:マルチエージェント運転のための現実的な都市環境

VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

シェア:XThreadsFacebookLINEはてブBluesky

LLM制御のエージェントが独立した目的を持ち、互いの行動が交通流やリスクに影響し合う都市運転ベンチマークを提案し、9モデルを評価した。

詳しい要約

1. どんなもの?

- 3D都市運転のベンチマーク VehicleArena を提案 - 独立した目的を持つ複数の embodied agents が同一環境で行動 - LLM 制御の agents が乗客要求を満たしつつ交通を走行 - 単一 agent と複数 agent の 112 評価タスクを提供 - 各 agent の運転判断が周囲の交通流・遅延・リスク・観測を変化させる

2. 先行研究と比べてどこがすごい?

- 既存 benchmark は共有目標や明示的な相互作用プロトコルを仮定 - 本研究は創発的な物理的結合 (emergent physical coupling) を扱う - 独立目的の agents が互いの条件を変える状況を評価可能 - 単一 agent だけでなく複数 agent の運転タスクを含む - 焦点車両以外への外部性 (externalities) を測定できる点が新しい

3. 技術・手法の肝は?

- LLM 制御の agents が進化する乗客要求を遂行 - 複雑な交通環境をナビゲートする 3D 都市運転シミュレータ - 各 agent の運転決定が交通流・遅延・リスク・後続観測を再形成 - 単一 agent と複数 agent の 112 評価タスクを設計 - シミュレータの native traffic controller と比較可能な評価設定

4. どうやって有効だと検証した?

- 9 つのモデルを評価 - 単一 agent タスクの最高到着率は 65.0% - 複数 agent タスクの最高到着率は 65.6% - 乗客要求や cabin スコアが高くても旅行完了に直結しないことを確認 - 一致した複数 agent 実行で、全 focal policy が周囲車両の到着率を native traffic controller より低下させることを確認

5. 議論はある?

- 高い乗客要求スコアや cabin スコアが到着率に信頼性高く結びつかない - 複数 agent 設定で焦点車両以外への外部性が測定された - 独立目的の agents 間の物理的結合が性能に影響する可能性 - 既存 benchmark の共有目標・明示プロトコル仮定の限界を示唆 - 具体的な議論の詳細は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明示されていない - 同分野の定番として multi-agent reinforcement learning (MARL) の運転ベンチマーク - 関連手法として LLM-controlled agents の embodied AI 研究 - 交通シミュレーションにおける native traffic controller の研究 - 外部性 (externalities) を扱う multi-agent 評価研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jie Yang, Jiajun Chen, Jiazheng Zhou, Mianqiu Huang, Yining Zheng, Yuxin Wang, Xipeng Qiu

分類: cs.MA, cs.CL, cs.CV

原文アブストラクト

Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator's native traffic controller, revealing measurable externalities beyond the focal vehicle itself.

PR本紙発行元 EmplifAI