日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
視覚グラウンディングarXiv:2609.27076

Pro-Bench: 実世界の異種環境におけるプロンプト頑健なオープンボキャブラリ視覚グラウンディングのベンチマーク

Pro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments

シェア:XThreadsFacebookLINEはてブBluesky

地下・産業・屋内・屋外・都市など多様な実環境のロボット画像13k枚以上と515の多様なクエリを用い、16のオープンボキャブラリモデルのゼロショット視覚グラウンディング性能とプロンプト頑健性を評価するベンチマークを提案した。

詳しい要約

1. どんなもの?

- 実世界の異種環境における open-vocabulary visual grounding のための prompt-conditioned benchmark である Pro-Bench を提案。 - 地下、産業、屋内、屋外、都市の独立したロボティクス領域から 13k+ の RGB フレーム、74.5k の手動 instance アノテーション、515 の target query を含む。 - query は categorical, attributive, relational, affordance, state, part-whole, negative, compositional semantics をカバー。 - 16 の open-vocabulary モデル構成を strict zero-shot 推論で評価し、IoU 閾値ごとの localisation 精度、end-to-end 推論遅延、prompt による性能変動、target recovery 一貫性を測定。

2. 先行研究と比べてどこがすごい?

- 既存 benchmark は短いカテゴリラベルと web-scraped 画像に大きく依存しており、実展開下での多様な query や視覚条件への頑健性が不明であった。 - Pro-Bench は実世界の異種ロボティクス領域から収集した RGB フレームと手動アノテーションを用い、多様な意味論的 query を網羅。 - prompt 条件付きで zero-shot 推論を評価し、prompt 頑健性や target recovery 一貫性を体系的に測定する点が新しい。

3. 技術・手法の肝は?

- 独立したロボティクス領域(subterranean, industrial, indoor, outdoor, urban)から RGB フレームを収集し、手動で instance アノテーションを付与。 - 515 の target query を categorical, attributive, relational, affordance, state, part-whole, negative, compositional semantics の観点で設計。 - 16 の open-vocabulary モデル構成を strict zero-shot で評価し、IoU 閾値ごとの localisation 精度、end-to-end 推論遅延、prompt による性能変動、target recovery 一貫性を指標化。

4. どうやって有効だと検証した?

- 16 の open-vocabulary モデル構成を strict zero-shot 推論でベンチマーク。 - IoU 閾値ごとの localisation 精度、end-to-end 推論遅延、prompt による性能変動、target recovery 一貫性を測定。 - 結果として、prompt 頑健性はアーキテクチャに強く依存し、10/16 のモデル構成が短いカテゴリラベルで最高性能、free-form query で最高精度は 1 つのみ。 - 同様の aggregate mAP でも reformulation 間の一貫した target recovery に大きな差が隠れうることを示した。

5. 議論はある?

- prompt 頑健性がアーキテクチャ依存であること、短いカテゴリラベルが多くのモデルで最良であること、free-form query が最良となるモデルは稀であることを指摘。 - aggregate mAP が reformulation 間の一貫した target recovery の差異を隠蔽しうることを議論。 - Pro-Bench がこれらのギャップの体系的評価と prompt-robust visual grounding の支援を可能にする。 - 具体的な限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている個別研究は明示されていない。 - 同分野の定番として open-vocabulary visual grounding の代表的手法(例: GLIP, Grounding DINO, OWL-ViT など)や、ロボティクス向け visual grounding benchmark を次に読む候補として挙げる。 - ただし要旨に具体的な論文名の記載はないため、これらは一般名としての提案である。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Linus Nwankwo, Muslim Alaran, Christian Rauch, Stanley Chukwuebuka Obilikpa, Elmar Rueckert

分類: cs.CV, cs.RO

原文アブストラクト

Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perceptual taxonomies. However, existing benchmarks largely rely on short category labels and web-scraped imagery, leaving it unclear whether open-vocabulary models can robustly ground diverse queries and visual conditions under real deployments. We introduce \textbf{Pro-Bench}, a prompt-conditioned benchmark for open-vocabulary visual grounding in heterogeneous, real-world environments. Pro-Bench includes $13k+$ RGB frames from independent robotic domains (subterranean, industrial, indoor, outdoor, urban), with $74.5k$ manual instance annotations and $515$ target queries covering categorical, attributive, relational, affordance, state, part-whole, negative, and compositional semantics. We benchmarked $16$ open-vocabulary model configurations in strict zero-shot inference, measuring localisation accuracy across IoU thresholds, end-to-end inference latency, prompt-induced performance variation, and target recovery consistency. Our results show that prompt-robustness is strongly architecture-dependent. Most model configurations ($10/16$) perform best with short category labels, whereas free-form queries yield the highest accuracy for only one. Moreover, similar aggregate mAP can conceal substantial differences in consistent target recovery across reformulations. Pro-Bench enables systematic evaluation of these gaps and supports prompt-robust visual grounding. Pro-Bench: https://pro-bench.github.io/.

関連論文

PR本紙発行元 EmplifAI