日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.19554

VA-Bench:視覚実演・能動的知覚・メートル制御による身体性空間知能の評価

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

シェア:XThreadsFacebookLINEはてブBluesky

RGB実演から手順を学び、カメラ視点を能動的に選び、メートル単位の動作指令を出してフィードバックで修正する「観察-推論-行動-修正」ループを評価するベンチマークVA-Benchを提案。最良モデルでもタスク成功率は約54%にとどまることを示した。

詳しい要約

1. どんなもの?

- 本論文は VA-Bench を提案する。 - 視覚デモンストレーションから手順的文脈を学び、能動的に視点を選び、metric な Cartesian コマンドを発行し、実行フィードバックで修正する observe-reason-act-revise ループを評価するベンチマーク。 - 14 の基本タスクファミリ(11 単腕、3 双腕)、7 の held-out な geometry/layout 変種、長 horizon の five-object composition track を含む。 - 12 の主要モデル条件を同一の 20 物理検証済み seed で 3 回独立実行し、terminal success、9 つの trajectory-level behavioral diagnostics、subtask progress を報告する。

2. 先行研究と比べてどこがすごい?

- 従来の spatial intelligence 評価は物体位置の記述に留まりがちだが、本ベンチマークは不完全観測下で不足証拠を同定・取得し、共通空間フレームで解釈し、行動に移す完全ループを測る点が新しい。 - モデルに privileged な object poses、oracle trajectories、learned action heads を与えず、RGB のみのデモンストレーションから手順的文脈を学ばせる。 - 固定の model-agnostic controller がモデル指定の target のみを実行する設定で評価する。 - 具体的な先行研究との比較は要旨からは不明。

3. 技術・手法の肝は?

- RGB のみのデモンストレーションから手順的文脈を学習。 - モデルが能動的に camera viewpoints を選択。 - metric な Cartesian commands を発行し、実行フィードバックから修正。 - 固定の model-agnostic controller がモデル指定の target のみを実行。 - 評価は terminal success、9 つの trajectory-level behavioral diagnostics、subtask progress で行う。

4. どうやって有効だと検証した?

- 12 の主要モデル条件を 3 回独立実行し、基本タスクごとに同一の 20 物理検証済み seed で評価。 - 最良モデルは annotated run で target localization 100.0%、spatial relations 78.9% だが、3 回の macro-average task success は 53.93±3.17%。 - 能動的 camera control は受動的多視点観察より有意に成功率を改善(一比較で 27.86% から 57.50% へ)。 - held-out な geometric transfer は成功率を 30 ポイント以上低下させ得る。 - 厳密な長 horizon エピソードを完了したモデルはなく、部分進捗は大きい。

5. 議論はある?

- 最良モデルでも target localization と spatial relations のスコアに比べ、macro-average task success は 53.93±3.17% と低い。 - 能動的 camera control の有効性が示される一方、held-out な geometric transfer で成功率が 30 ポイント以上低下する。 - 厳密な長 horizon エピソードを完了するモデルは存在せず、部分進捗は大きい。 - 一般-purpose MLLMs が視覚デモと能動取得証拠を embodied action に変換できるかを問う。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明示されていない。 - 同分野の定番として、embodied AI の vision-language-action モデル、active perception、spatial reasoning ベンチマーク(例: RLBench、ALFRED、Habitat、BEHAVIOR)が次に読む候補となる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu

分類: cs.RO, cs.CV

原文アブストラクト

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.

関連論文

PR本紙発行元 EmplifAI