日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音響知覚arXiv:2609.23407

OmniEcho: 身体性エージェントのための空間音響理解

OmniEcho: Spatial Audio Understanding for Embodied Agents

シェア:XThreadsFacebookLINEはてブBluesky

実世界の空間音響・視覚シーンを集めたベンチマークOmniEchoBenchを構築し、一人称アンビソニックス音響を扱う全モーダルモデルOmniEchoを提案。音響による知覚・ナビゲーションで最先端性能を達成した。

著者: Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong

分類: cs.SD, cs.AI

原文アブストラクト

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce \textbf{OmniEchoBench}, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments. To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose \textbf{OmniEcho}, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation. These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges.

関連論文

PR本紙発行元 EmplifAI