MECoBench: 具現化環境におけるマルチモーダルエージェント協調の体系的研究
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
マルチモーダル大規模言語モデルを具現化エージェントとして協調させるためのベンチマークMECoBenchを提案し、多様な実世界タスクでの協調効果と限界を実験的に分析した。
著者: Qingyun Liu, Jiwen Zhang, Jingyi Hu, Siyuan Wang, Zhongyu Wei
分類: cs.MA, cs.AI, cs.CL, cs.CV
原文アブストラクト
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBench, a multimodal embodied cooperation benchmark with an evaluation platform spanning diverse real-world tasks, two cooperation structures, and three collaboration modes. Through extensive experiments across various MLLMs, we summarize three key findings: (i) Collaboration generally improves embodied task completion, but its benefits depend on balancing collaborative gains against coordination complexity. (ii) Communication is essential to collaboration gains, while the best collaboration mode depends on team size and model capability. (iii) Moreover, collaboration improves robustness under noisy priors and exploration conditions. Generally, MECoBench provides a systematic testbed for understanding the mechanisms and limits of multimodal embodied collaboration. Code and dataset are available at https://github.com/q-i-n-g/MECoBench.
関連論文
- SyncPlan: 明示的同期と適応的修正による長期的LLM協調マルチエージェント協調
- LLawCo: 協調の法則を学習して具現化マルチエージェント行動をモデル化するマルチエージェント協調