日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーションarXiv:2610.06510

MarvisNav: ゼロショット物体ナビゲーションにおける経路選択のための記憶の可視化

MarvisNav: Making Memory Visible on Route Choices for Zero-Shot Object Navigation

シェア:XThreadsFacebookLINEはてブBluesky

探索履歴を視覚的な経路候補に直接重ねて表示することで、VLMが目標関連性と探索状態を同時に評価できるゼロショット物体ナビゲーション手法を提案し、HM3Dで最高性能を達成した。

詳しい要約

1. どんなもの?

- ゼロショット物体ナビゲーション(ZSON)のためのフレームワーク「MarvisNav」を提案。 - 探索履歴を視覚的な経路選択肢上に直接可視化する。 - トポロジカルグラフを維持し、候補ノードとその探索状態を自己中心視点に投影。 - 視覚言語モデル(VLM)が目標関連性と探索状態を同時に評価可能。 - ポリシー学習なしでHM3Dで81.2% SR、42.5% SPLを達成。

2. 先行研究と比べてどこがすごい?

- 従来のZSONではVLMが自己中心画像から探索領域を推論し、探索履歴はテキストやマップで別表現。 - そのため融合ステップが必要か、記憶と経路選択の対応が暗黙的。 - MarvisNavは探索記憶を視覚的経路選択肢上に直接表示し、事後融合や再ランキングを不要に。 - ポリシー学習なしでHM3Dで最先端性能、MP3Dでも競争力。 - VLM呼び出し回数がWMNavの7.5%と大幅に少ない。

3. 技術・手法の肝は?

- トポロジカルグラフを維持し、候補ノードを自己中心視点に投影。 - 各候補ノードに探索状態(二値的な訪問だけでなく局所的な探索進捗)を付与。 - 探索状態を視覚的候補に直接結びつけることで、VLMが目標関連性と探索状態を統合評価。 - 別途の融合や再ランキング段階を排除。 - ポリシー学習は不要。

4. どうやって有効だと検証した?

- HM3Dデータセットで81.2% SR、42.5% SPLを達成し最先端性能を確認。 - MP3Dデータセットでも競争力のある性能を示す。 - 代表的なVLMベース手法と比較し、VLM呼び出し回数がWMNavの7.5%と大幅に少ない。 - 多様なシーンでの実ロボット実験により実用性を検証。

5. 議論はある?

- 記憶表現がVLMの決定とZSON性能を形成することを示す。 - 効果的な記憶利用はその可用性だけでなく、表現方法にも依存することを強調。 - 具体的な議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- WMNav(VLMベースの代表手法として比較) - その他のVLMベースZSON手法(具体的名称は要旨からは不明) - トポロジカルグラフを用いたナビゲーション手法 - ゼロショット物体ナビゲーション(ZSON)の一般的な関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jincheng Wang, Chi Pui Chan, Wei Zeng, Shuyang Zhang, Jianhao Jiao, Dimitrios Kanoulas

分類: cs.RO

原文アブストラクト

When searching for an object, people choose their next move by considering both likely target locations and places already explored. The current view can cue place-associated memories, bringing target relevance and prior exploration into the same spatial context. In many zero-shot object navigation (ZSON) methods, however, vision-language models (VLMs) infer promising search areas from egocentric images, while exploration history is represented separately, e.g., as text or maps. This separation either requires an additional fusion step or leaves the correspondence between memory and route choices implicit for the VLM to recover. We instead make exploration memory directly visible on visual route choices. We propose MarvisNav, a ZSON framework that maintains a topological graph and projects candidate nodes together with their exploration states onto egocentric views as memory-bearing visual route choices. These states capture local exploration progress beyond binary visitation. By binding exploration state directly to each visual candidate, MarvisNav enables the VLM to jointly evaluate target relevance and exploration state without a separate post-hoc fusion or reranking stage. Without policy training, MarvisNav achieves state-of-the-art performance on HM3D (81.2% SR and 42.5% SPL), while remaining competitive on MP3D. It also outperforms representative VLM-based methods with far fewer VLM calls (e.g., 7.5% of WMNav). Real-robot experiments across diverse scenes further validate its practical deployability. Beyond MarvisNav, our study shows that memory representation shapes VLM decisions and ZSON performance, highlighting that effective memory use depends not only on its availability, but also on how it is represented. Code and project page will be available at \url{https://wangjincheng1998.github.io/MarvisNav/}.

関連論文

PR本紙発行元 EmplifAI