IMPRINT: 画像条件付きクエリ拡張によるロングテール物体ゴールナビゲーション
IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
テキストのみのクエリに依存する既存の物体ゴールナビゲーション手法の限界を克服するため、ウェブから取得した画像でクエリを拡張し、視覚言語モデルを用いて意味マップと照合することで物体の位置特定精度を向上させるゼロショットフレームワークを提案した。
著者: Jelin Raphael Akkara, Filippo Ziliotto, Luciano Serafini, Lamberto Ballan, Tommaso Campari
分類: cs.CV
原文アブストラクト
Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing approaches typically depend on text-only queries, which become less reliable as semantic specificity increases toward fine-grained object categories. We introduce IMPRINT, a zero-shot plug-and-play framework that enriches textual object queries with web-sourced images to improve grounding in queryable maps. Retrieved images are encoded using a vision-language model, matched against the semantic map to produce similarity maps, and aggregated to yield context-aware localization. Notably, this requires no training or modification of the underlying navigation policy. To explicitly evaluate long-tail behavior, we present HSSD-rare, a new ObjectNav benchmark built on Habitat Synthetic Scenes and featuring semantically specific subcategories. Across both OVON and HSSD-rare, image-conditioned queries consistently improve object grounding and yield end-to-end navigation gains. Further analysis reveals that translating localization gains to navigation performance depends critically on downstream detection quality, highlighting a key systems bottleneck in long-tail embodied navigation.