日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10388

RoboQuest:探索・検査・試験を行う汎用物理エージェント

RoboQuest: Generalist Physical Agents that Search, Inspect and Test

シェア:XThreadsFacebookLINEはてブBluesky

未知環境で必要な情報を物理的相互作用を通じて能動的に獲得し、タスク完了を自律的に判断する汎用エージェントのベンチマークRoboQuestを提案。最先端のマルチモーダルエージェントでも成功率は23%と低く、失敗の多くは探索の早期終了に起因することを示した。

詳しい要約

1. どんなもの?

- 目的志向の embodied exploration のための benchmark「RoboQuest」を提案 - 10 個の mobile manipulation タスクで構成 - 3 種類の不確実性に焦点: search、manipulation-based inspection、interactive testing - agent は物理的相互作用を通じて task-relevant information を能動的に獲得し、その evidence で subsequent actions を適応させ、自律的に task completion への commit を判断する必要がある - 5 つの frontier multimodal agents と full-episode demonstrations で fine-tune した $\pi_{0.5}$ policy を共通の visuomotor interface で評価

2. 先行研究と比べてどこがすごい?

- 従来の multimodal foundation models は manipulation タスクの generalist physical agents として有望だが、観測に無い情報を interaction で獲得する能力は未検証 - RoboQuest は search、inspection、testing という 3 形態の不確実性を明示的に扱う初の benchmark - 単なる execution 能力ではなく、情報獲得・適応・commit 判断という goal-directed embodied exploration を評価 - full-episode demonstrations を公開し、fine-tuned policy の baseline も提供

3. 技術・手法の肝は?

- 10 個の mobile manipulation タスクを search、manipulation-based inspection、interactive testing の 3 カテゴリに分類 - 共通の visuomotor interface で 5 つの frontier multimodal agents と $\pi_{0.5}$ policy を評価 - full-episode demonstrations を収集・公開し、$\pi_{0.5}$ を fine-tune - 実行スキルの isolated tests を実施し、隠された情報を与えた場合の動作可否を検証 - failure analysis により、失敗要因を execution と exploration の早期停止などに分類

4. どうやって有効だと検証した?

- 5 つの frontier multimodal agents と fine-tuned $\pi_{0.5}$ policy を RoboQuest の 10 タスクで評価 - 最良の agent でも成功率は 23%、fine-tuned policy はほぼ成功しない - 隠された情報を与えた isolated execution skill tests では、agents は要求される動作のほとんどを実行可能 - failure analysis により、失敗の少数のみが execution に起因し、多くは exploration の早期停止や disturbance の未修復、trial and error 学習の困難さに起因すると判明

5. 議論はある?

- agents は必要な evidence を観測する前に意思決定し、exploration を早期に停止する傾向 - agents は自身の exploration による disturbance を防止・修復することがほとんどない - ほとんどのモデルで trial and error による学習が依然として困難 - execution 能力は isolated tests で示されるため、主なボトルネックは exploration と情報に基づく適応・commit 判断 - 要旨からは、タスク設計や評価指標の詳細、モデル間の具体的な比較は不明

6. 次に読むべき論文は?

- $\pi_{0.5}$ policy(fine-tune 元の手法) - multimodal foundation models を用いた generalist physical agents の先行研究 - embodied exploration や active information acquisition に関する研究 - mobile manipulation の benchmark(例: 一般的な manipulation benchmark) - trial and error 学習や disturbance recovery に関する研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a $π_{0.5}$ policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.

関連論文

PR本紙発行元 EmplifAI