日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.35032

JRDB-AVR: 実世界環境における身体性エージェントのための能動的視覚推論ベンチマーク

JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments

シェア:XThreadsFacebookLINEはてブBluesky

実世界のロボティクスデータから、時間・視点を指定して観測を要求しながら視覚推論を行う能動的エージェント向けベンチマークを構築し、回答だけでなく根拠となる視覚証拠も評価する。

詳しい要約

1. どんなもの?

- 実世界のembodied visual reasoning向けbenchmark - JRDBのロボティクスデータから構築 - 質問生成engineで作成 - エージェントは視野が限定 - 証拠が時間・視点・対象間に分散 - 質問に対しtimestampとviewing angleで観測要求 - 最終回答と根拠visual evidenceの両方を評価 - 多様な質問を含む - temporal search - viewpoint selection - human-oriented compositional reasoning - 参照手法JRDB-AVR-Agentも提案

2. 先行研究と比べてどこがすごい?

- 従来のvisual reasoning benchmarkは受動的観測と最終回答を評価 - 能動的reasoningとevidence獲得を見落とす - 本benchmarkは能動的evidence-aware評価を明示 - 回答だけでなく根拠visual evidenceも評価 - 現行baselineで回答精度とevidence精度に大きな乖離 - VLMが根拠のない正解を出しうることを示す - embodied visual reasoningには能動的evidence-aware評価が必要と主張

3. 技術・手法の肝は?

- JRDBの実世界データを利用 - 構造化question-generation engineで質問を生成 - 時間・視点・対象をまたぐ証拠を要求 - エージェントはtimestampとviewing angleを指定しbounded observationsを要求 - 評価は最終回答とgrounded visual evidenceの両方 - 参照手法JRDB-AVR-Agent - 明示的なobservation-grounded graph-based world modelを維持 - solvingを通じて回答

4. どうやって有効だと検証した?

- JRDB-AVR上でbaselineを評価 - 回答精度とevidence精度の乖離を実験で確認 - 現行VLMがunsupported correct answersを出しうる - 複数実環境・多様な質問で検証 - temporal search - viewpoint selection - human-oriented compositional reasoning - 詳細な実験設定・指標は要旨からは不明

5. 議論はある?

- 回答精度とevidence精度の乖離が主要な議論 - 正解でも根拠が支持されない場合がある - 能動的evidence-aware評価の必要性を主張 - 限定的視野・分散証拠という設定の重要性 - 具体的な限界・失敗事例・倫理的議論は要旨からは不明

6. 次に読むべき論文は?

- JRDB-AVR-Agent(本論文の参照手法) - JRDB(元データセット) - 関連するvisual reasoning benchmark - 能動的観測やevidence評価を扱うもの - embodied visual reasoningの定番手法 - Vision-Language Models (VLMs) - graph-based world model - 具体的な参照論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhixi Cai, Fucai Ke, Sukai Huang, Maria Garcia de la Banda, Peter J. Stuckey, Gholamreza Haffari, Hamid Rezatofighi

分類: cs.AI, cs.CV, cs.RO

原文アブストラクト

In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active reasoning and evidence acquisition. We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning agentic method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning. Code and benchmark are available at https://github.com/ControlNet/JRDB-AVR.

関連論文

PR本紙発行元 EmplifAI