日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
空間音声理解arXiv:2609.34347

SAIL: 分離音響・空間エンコーディングとデュアルストリームQ-Formerによる大規模言語モデルを用いた空間音声知能

SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former

シェア:XThreadsFacebookLINEはてブBluesky

音響特徴と空間特徴を別々のストリームで扱い、音源ごとの対応を保ちながらLLMに接続する空間音声理解フレームワークSAILを提案し、複数音源のイベント検出・方向・距離推定・空間推論を改善した。

著者: Zhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou, Woon-Seng Gan, Wenwu Wang

分類: cs.SD, cs.AI, eess.AS

原文アブストラクト

Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.

PR本紙発行元 EmplifAI