日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ベンチマークarXiv:2610.10409

RobotWorld:多様なタスクと身体性にわたるロボット利用のためのマルチモーダルエージェントベンチマーク

RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

シェア:XThreadsFacebookLINEはてブBluesky

汎用エージェントがロボットインターフェースを通じて物理タスクを実行できるかを評価する84タスクのシミュレーションベンチマークを提案し、能力の転移と失敗要因を分析した。

詳しい要約

1. どんなもの?

- 汎用エージェントの物理世界での能力を評価するためのシミュレーションテストベッド「RobotWorld」を提案。 - 84のタスクを含み、manipulation, mobile manipulation, locomotion, driving, aerial controlの5領域をカバー。 - 各タスクには明示的なinteraction budgetsと実行可能なsuccess checksが設定されている。 - 指示と観測をロボットインターフェース経由で物理的タスク実行に変換する能力を測定。 - 実行トレースとタスク結果を分析し、転移可能な能力と信頼性を妨げるギャップを特定。

2. 先行研究と比べてどこがすごい?

- 従来のデジタルタスク向けベンチマークと異なり、物理世界でのロボット利用に焦点を当てた点が新しい。 - 多様なembodimentsとタスク領域を横断する統一的な評価環境を提供。 - 単なる成功率だけでなく、実行トレースを分析して能力の転移と失敗要因を明らかにする点が特徴。 - 既存研究では明示されていなかった、モデル間の成功傾向の違い(例:Astra vs Opus 5.5)を実証的に示した。

3. 技術・手法の肝は?

- シミュレーション上でロボットインターフェースを介し、指示と観測から物理タスクを実行するエージェントを評価。 - タスクごとにinteraction budgetsとsuccess checksを定義し、定量的に成否を判定。 - 実行トレースを収集し、知覚・制御ワークフローの構築能力(image segmentation, camera calibration, spatial estimation, dynamics-based computation)を分析。 - タスク結果と実行行動を結びつけ、能力の転移パターンと失敗モードを特定。

4. どうやって有効だと検証した?

- 84タスクにわたるエージェントの実行結果とトレースを分析。 - 現在のエージェントが高度な知覚・制御ワークフローを構築できることを確認。 - しかし、それらが一貫して成功行動に結びつかないことを実証(例:指示されたポーズに到達しても物体状態を失う、無効な行動を修正できない、回復が遅れる、未完了タスクを完了と誤認)。 - モデル間の成功傾向の違い(Astraは空間・制約接触目標、Opus 5.5は連続バランス・時間制約相互作用目標で成功が多い)を定量的に示した。

5. 議論はある?

- 汎用エージェントの能力は物理世界に均一に転移しないという課題を提示。 - 知覚・制御ワークフローの構築能力と、実際の成功行動との間にギャップがあることを議論。 - 失敗モードとして、物体状態の喪失、無効行動の未修正、回復の遅れ、未完了タスクの誤認を挙げる。 - モデルごとの成功傾向の違いが、トレーニングや設計のターゲットになり得ると示唆。 - より信頼性の高い物理世界エージェントのための具体的な目標を設定。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、embodied AI, robot learning, multimodal agents, simulation benchmarks (例: RLBench, ALFRED, Habitat) などが同分野の定番として挙げられる。 - 具体的な次読論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu, Weidong Huang, Jiankai Sun, Haitao Li, Zijian Wu, Yuzhi Huang, Fanding Huang, Hanwen Sun, Jiashun Liu, Jingqi Tong, Mingxin Huang, Shaoli Hu, Shijue Huang, Tianyi Bai, Xinyuan Wang, Yunlong Lin, Zhengyang Tang, Zhexin Zhang, Zhuo Chen, Xierui Song, Juntao Dai, Boyuan Chen, Jiaming Ji, Fangneng Zhan, Mengkang Hu, Wei Xue, Yonggang Zhang, Han Hu, Tsung-Yi Ho, Yike Guo

分類: cs.RO, cs.LG

原文アブストラクト

General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.

関連論文

PR本紙発行元 EmplifAI