日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
空間音声理解arXiv:2610.05610

SEA-LM: ウェアラブルマイクアレイのための自己中心空間音声理解

SEA-LM: Egocentric Spatial Audio Understanding for Wearable Microphone Arrays

シェア:XThreadsFacebookLINEはてブBluesky

スマートグラス型マイクアレイの音声から空間情報を捉え、音源定位や複数話者の選択的書き起こしを行うマルチモーダル大規模言語モデルを提案した論文。

著者: Sonal Kumar, Sinan Hersek, Artem Dementyev, Mengzhen Pan, Ishan Chatterjee, Anurag Kumar, Ramani Duraiswami, Dinesh Manocha, Andrea Colaco

分類: cs.SD, cs.AI

原文アブストラクト

Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.

関連論文

PR本紙発行元 EmplifAI