RAO-Nav: 統合言語モデルによるゼロショット音声視覚ナビゲーションの探求
RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation
音声と視覚を統合した大規模言語モデルをゼロショットの音声視覚ナビゲーションに適用し、学習データなしで既存手法を上回る性能を実現した。
著者: Qilang Ye, Meng Liu, Yu Zhou
分類: cs.AI
原文アブストラクト
We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: https://github.com/rikeilong/OmniAV\_Nav.