日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.26360

階層的間取り誘導型視覚言語探索による身体性質問応答

Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering

シェア:XThreadsFacebookLINEはてブBluesky

RGB-D観測から階層的シーングラフと開放語彙占有マップを構築し、間取りの事前情報とVLM計画を組み合わせて未知環境を効率的に探索するEQAフレームワークを提案した。

詳しい要約

1. どんなもの?

- Embodied Question Answering (EQA) のためのフレームワーク HFLEX-EQA を提案 - 未見環境を探索し、シーンに関する質問に答えるエージェントを対象 - オンラインでの scene graph 構築、VLM ベースの計画、semantic frontier exploration、floorplan priors を階層的に統合 - RGB-D 観測から階層的 scene graph と open-vocabulary occupancy map を逐次構築 - VLM が scene graph、タスク関連視覚観測、探索履歴、推定 topological floorplan を統合推論 - room-discovery 戦略により、意味的に関連する未観測の部屋タイプへ探索を誘導

2. 先行研究と比べてどこがすごい?

- 近年の VLM と semantic map / scene graph を用いる EQA 手法は、局所観測のみで探索を駆動し、環境の構造的 prior をほとんど活用していない - 本研究は floorplan priors を探索に組み込み、構造的 prior を活用する点が新しい - VLM ベースの階層的計画と構造的 floorplan prior の組み合わせが EQA タスクに有効であることを示す - 具体的な先行研究名や定量的な比較優位は要旨からは不明

3. 技術・手法の肝は?

- オンラインでの階層的 scene graph 構築と open-vocabulary occupancy map の逐次生成 - VLM が scene graph、タスク関連視覚観測、探索履歴、推定 topological floorplan を統合して推論 - semantic frontier exploration を採用 - room-discovery 戦略:floorplan と open-vocabulary frontier semantics を利用し、意味的に関連する未観測の部屋タイプへ探索を誘導 - 階層的な計画と構造的 prior を組み合わせた探索制御

4. どうやって有効だと検証した?

- OpenEQA と ExploreEQA ベンチマークで評価 - 実屋内環境で quadruped robot に実装して展開 - VLM ベースの階層的計画と構造的 floorplan prior の組み合わせの利点を実証 - 具体的な評価指標や数値結果は要旨からは不明

5. 議論はある?

- 要旨からは不明 - 限界や失敗事例、計算コスト、実環境での課題などについての議論は要旨に記載なし

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として Vision-Language Models (VLMs)、semantic maps、scene graphs を用いた EQA アプローチ - ベンチマークとして OpenEQA、ExploreEQA - 同分野の定番として Embodied Question Answering (EQA) の一般的な手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Albert Gassol Puigjaner, Kostas Alexis

分類: cs.RO

原文アブストラクト

Embodied Question Answering (EQA) requires an agent to explore a previously unseen environment, gather relevant information, and answer questions about the scene. Recent approaches leverage Vision-Language Models (VLMs) together with semantic maps or scene graphs to guide exploration. However, exploration is typically driven only by local observations, while structural priors about the environment remain largely unused. We propose HFLEX-EQA, a hierarchical EQA framework that combines online scene graph construction, VLM- based planning, semantic frontier exploration, and floorplan priors. The system incrementally builds a hierarchical scene graph and an open-vocabulary occupancy map from RGB-D observations, enabling a VLM to jointly reason over the scene graph, task-relevant visual observations, exploration history, and an estimated topological floorplan. Furthermore, we introduce a room-discovery strategy that leverages the floorplan and open-vocabulary frontier semantics to guide exploration toward semantically relevant yet currently unobserved room types. We evaluate HFLEX-EQA on the OpenEQA and ExploreEQA benchmarks and demonstrate deployment on a quadruped robot in real indoor environments. Our results demonstrate the benefit of combining VLM-based hierarchical planning with structural floorplan priors for the EQA task.

関連論文

PR本紙発行元 EmplifAI