日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00854

フロンティアVLMエージェントはロボット汎用istになれるか?Embodied Agent Arenaによる実証研究

Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena

シェア:XThreadsFacebookLINEはてブBluesky

VLMのロボット汎用性を評価するベンチマーク「Embodied Agent Arena」を提案し、7つのVLMを幾何推定・空間推論・アフォーダンス・タスク計画・操作の観点で比較した。精密推定は得意だが、協調的な目標指向行動の完遂に課題が残ることを示した。

詳しい要約

1. どんなもの?

- Frontier VLM を robot generalist として評価するための Embodied Agent Arena を提案 - Geometry, Spatial Reasoning, Affordance, Task Planning, Manipulation の5領域を横断 - 32 の既存ソースと新規 benchmark GeoProbe から計 1,000 ケースを収録 - GeoProbe は Blender renders と real-scene images 上の geometric estimation 用 - 7 つの VLM を評価し Astra の task-specific な優位性を分析

2. 先行研究と比べてどこがすごい?

- 局所的な能力評価に留まらず、complete task success への寄与を測る点が新しい - 単一タスクでなく Geometry から Manipulation まで統合的に横断 - 32 の established sources を集約しつつ GeoProbe を追加 - minimal harness で source observations と operations を保持 - metric precision, functional grounding, native goal completion を分離して評価

3. 技術・手法の肝は?

- Embodied Agent Arena と minimal harness が中核 - source observations と operations を保ちつつ評価軸を分離 - metric precision, functional grounding, native goal completion を個別に測定 - richer-observation execution protocols と multi-round review を比較 - GeoProbe で Blender renders と real-scene images の geometric estimation を評価

4. どうやって有効だと検証した?

- 7 つの VLM を Arena 全体で評価 - Astra の task-specific advantages を分析 - richer-observation execution protocols と multi-round review を比較 - Astra の優位は precise estimation と usable-contact localization で最大 - coordinated, goal-directed actions の完遂が依然として主要な gap と確認

5. 議論はある?

- 局所能力が complete task success に繋がるか、どこで不足するかを議論 - Astra は precise estimation と usable-contact localization に強い - 一方で coordinated, goal-directed actions の完遂が robot generalism への鍵 - richer-observation execution protocols と multi-round review の効果を比較 - 具体的な議論の詳細は要旨からは不明

6. 次に読むべき論文は?

- GeoProbe - Embodied Agent Arena 内の 32 established sources - Astra - richer-observation execution protocols - multi-round review

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haojian Huang, Pukun Zhao, Zexi Li, Yehang Zhang, Yangkai Wei, Wenqian Li, Han Yang, Kaiwen Zhou, Ying-Cong Chen, Yinchuan Li

分類: cs.RO

原文アブストラクト

Frontier vision-language models (VLMs) combine scene estimation, interaction grounding, and executable actions. Understanding how these abilities support complete robotic tasks is central to evaluating their readiness as robot generalists. We introduce Embodied Agent Arena to examine where local competence supports, or falls short of, complete task success across Geometry, Spatial Reasoning, Affordance, Task Planning, and Manipulation. The arena contains 1,000 cases drawn from 32 established sources and GeoProbe, our new benchmark for geometric estimation on Blender renders and real-scene images. A minimal harness preserves source observations and operations while separating metric precision, functional grounding, and native goal completion. We evaluate seven VLMs, analyze Astra's task-specific advantages, and compare richer-observation execution protocols and multi-round review. Across the arena, Astra's advantage is strongest in precise estimation and usable-contact localization; completing coordinated, goal-directed actions remains the key gap to robot generalism.

関連論文

PR本紙発行元 EmplifAI