日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音源ナビゲーションarXiv:2609.37084

BCNav: 方位条件付き深度方策による音源ナビゲーション

BCNav: Bearing-Conditioned Depth Policies for Sound Source Navigation

シェア:XThreadsFacebookLINEはてブBluesky

音源方向推定(DOA)と深度画像のみを入力とするナビゲーション方策を模倣学習で訓練し、音響シミュレーションや実環境での音声データ収集なしに未知環境で音源へ移動できるロボットを実現した。

著者: Yaozhong Kang, Jiang Wang, Takeshi Ashizawa, Benjamin Yen, Kazuhiro Nakadai

分類: cs.RO

原文アブストラクト

The ability to navigate toward sound sources extends a robot's reach beyond its visual field, enabling response to auditory events in unknown environments. To equip robots with this capability, existing methods couple acoustic and visual information through joint audio-visual learning in acoustic simulators. However, acoustic simulation is both low-fidelity and expensive, producing a domain gap that prevents reliable real-world deployment, while the discrete action spaces inherited from grid-based simulators introduce an additional kinematic gap on physical robots. To alleviate these issues, we propose BCNav, a decoupled framework that separates the acoustic module from the learned navigation policy using direction-of-arrival (DOA) estimation: an estimator provides a scalar bearing to the sound source, so the navigation policy only processes depth images and a bearing angle, two inputs whose domain gaps are well characterized. We collect shortest-path demonstrations with calibrated bearing noise injection and train the policy via imitation learning to output continuous velocity commands directly executable on ground robots. We demonstrate the method in simulation and on a physical robot, navigating unknown environments without any acoustic fine-tuning, prior mapping, or real-world audio data collection. Code is available at https://github.com/york1to/bcnav.

PR本紙発行元 EmplifAI