日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLA/空間推論arXiv:2609.06880

文脈的観察者接地:視覚言語モデルにおける状況的空間推論の評価

Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models

シェア:XThreadsFacebookLINEはてブBluesky

ロボットなどの身体化タスクで、話者の視点に基づく空間関係の理解を評価するベンチマークPOVBenchを構築し、最新の視覚言語モデルが文脈から観察者の視点を推論して対象を特定できるかを検証した。

著者: Mimo Shirasaka, Haochen Zhang, Yonatan Bisk

分類: cs.CV, cs.CL, cs.LG, cs.RO

原文アブストラクト

Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated perspective. Humans infer such perspectives from shared environmental knowledge, activity context, and commonsense. Recent vision-language models (VLMs) appear capable of spatial reasoning, but their ability to infer a speaker's viewpoint from contextual cues and interpret situated spatial relations from that viewpoint remains unclear. We call this capability contextual observer grounding. To study this capability, we construct the Point-of-View Benchmark (POVBench), a dataset of 3D scenes and queries that disentangles Inferred, Stated, and Given forms of observer grounding in natural embodied communication. Given multi-view observations and a natural-language sentence, models must localize unseen or underspecified targets from situated spatial and contextual cues. Across state-of-the-art VLMs, localizing targets from directional language remains challenging, even when observer grounding is made explicit. We find that explicit breakdowns of observer-relative spatial reasoning improve target localization. Our project page is available at https://mimo-owl.github.io/POVBench/.

関連論文