日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34384

RoboIRGBench:視覚言語行動モデルにおける暗黙的指示対象接地のベンチマーク

RoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

ロボット操作における視覚言語行動モデルが、指示に明示されない暗黙的な対象・数量・関係を文脈から推論する能力を評価するベンチマークRoboIRG-Benchを提案し、既存モデルに顕著な性能低下があることを示した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルにおける Implicit Referential Grounding (IRG) 能力を評価する manipulation benchmark『RoboIRG-Bench』を提案。 - 既存 benchmark は指示にタスク関連情報が明示されていると仮定するが、実世界では対象・数量・関係が暗黙的に言及される。 - RoboMME を基盤に、11 タスクから派生した 40 変種を含み、direct・reasoning-mediated・spatial・contextual の 4 課題をカバー。 - 異なる memory 機構を持つ代表的な VLA を評価し、referential robustness gap を明らかにする。

2. 先行研究と比べてどこがすごい?

- 既存 benchmark は明示的指示を前提とし、暗黙的参照の回復能力を体系的に評価していない。 - 本研究は IRG を独立した未探索能力として位置づけ、初めて体系的 benchmark を提供。 - 明示的指示で高性能なモデルが、同じ情報を context から回復する必要がある場合に急激に性能低下することを示す。 - reasoning-mediated と spatial 参照が特に困難であることを明らかにし、外部 VLM 使用モデルでも失敗が残ることを示す。

3. 技術・手法の肝は?

- RoboMME を基に 11 タスクから 40 変種を構築し、direct・reasoning-mediated・spatial・contextual の 4 参照課題を設計。 - IRG は以前に確立された context の保持と検索を要することが多いため、異なる memory 機構を持つ VLA を評価対象とする。 - 外部 VLM を用いるモデルとそうでないモデルを比較し、外部 VLM をより強力なモデルに置き換えても gap が解消しないことを検証。 - Franka Research 3 ロボットアーム上で実世界 manipulation でも検証。

4. どうやって有効だと検証した?

- RoboIRG-Bench 上で代表的な VLA を評価し、referential robustness gap を定量的に確認。 - reasoning-mediated と spatial 参照が特に困難であること、外部 VLM 使用モデルがより robust だが依然重大な失敗を示すことを報告。 - 外部 VLM をより強力なモデルに置き換えても gap が残ることを確認。 - Franka Research 3 ロボットアームで実世界 manipulation を行い、gap が誤った referent grounding と下流の実行失敗として現れることを検証。

5. 議論はある?

- IRG は信頼性の高いロボット指示追従にとって独立かつ未探索の能力であると主張。 - 言語・知覚・推論・行動を robust に統合できる VLA の必要性を強調。 - 外部 VLM の強化だけでは gap が解消しないことを示し、モデル設計上の課題を提起。 - 具体的な限界や今後の課題の詳細は要旨からは不明。

6. 次に読むべき論文は?

- RoboMME(本 benchmark の基盤) - 代表的な VLA モデル(要旨では具体名が挙げられていないため一般名で記載) - 外部 VLM を用いる VLA モデル - Franka Research 3 を用いた実世界 manipulation 研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Aernaer Akelijiang, Jiannan Li, Zhineng Chen, Jingjing Chen, Bin Zhu

分類: cs.RO, cs.CV

原文アブストラクト

Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicitly, requiring robots to recover the intended target from linguistic and perceptual context. We study this capability as Implicit Referential Grounding (IRG) and introduce RoboIRG-Bench, a manipulation benchmark designed to systematically evaluate it. Built upon RoboMME, RoboIRG-Bench contains 40 variants derived from 11 tasks and covers four challenges, including direct, reasoning-mediated, spatial, and contextual referential grounding. As IRG often requires retaining and retrieving previously established context, we evaluate representative VLAs spanning different memory mechanisms. Our evaluation reveals a noticeable referential robustness gap. Models that perform well under explicit instructions can degrade sharply when the same task-relevant information must be recovered from context. Reasoning-mediated and spatial references are particularly challenging, while models using external VLMs show greater robustness but still exhibit significant failures. Moreover, replacing the external VLM with a stronger model does not eliminate these gaps. We further validate these findings on a Franka Research 3 robot arm, where the gap persists under real-world manipulation and manifests as both incorrect referent grounding and downstream execution failures. These results establish IRG as a distinct and underexplored capability for reliable robotic instruction following and highlight the need for VLAs that can robustly integrate language, perception, reasoning, and action.

関連論文

PR本紙発行元 EmplifAI