日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音響/ワールドモデルarXiv:2609.06837

BinauralVAE: ワールドモデルのための空間音響再構成

BinauralVAE: Spatial Audio Reconstruction For World Models

シェア:XThreadsFacebookLINEはてブBluesky

視覚に依存しがちな既存のワールドモデルに対し、空間音響に着目し、ロボットのナビゲーション行動と音響結果の因果関係を学習するためのパイプラインを提案。様々なVAEアーキテクチャを評価し、バイノーラル信号の潜在表現を獲得する。

著者: Luis Vitor Zerkowski, Luiz Velho

分類: cs.SD, cs.LG

原文アブストラクト

Embodied artificial intelligence has historically very much relied on visual perception, leading to a proliferation of multiple vision-centric world models. However, this reliance fails to capture spatial understanding in its entirety and can even present vulnerabilities in environments with visual occlusions, low-light conditions, or blackouts-scenarios, where acoustic information becomes a critical alternative for spatial awareness and navigation. Despite its potential, research into realistic spatial audio and particularly the development of audio-centric world models remains sparse. In this technical report, we introduce BinauralVAE: a flexible, open-source pipeline (https://github.com/Luizerko/BinauralVAE) that explores multiple models for spatialized audio reconstruction, progressing from fundamental baselines to advanced, mathematically grounded architectures. Our approach evaluates various Variational Autoencoder architectures -- including complex-valued variants -- to learn robust latent representations of binaural signals. Developed alongside AudioWorldSim, our methodology leverages realistic acoustic data captured as a simulated robot navigates an environment. This pipeline establishes a foundation for state representation in a future audio-based world model, designed to map the direct causal connection between navigational actions and their resulting acoustic consequences, and helping to enable sound as an essential complementary modality for spatial knowledge acquisition.