AudioWorldSim: 世界モデルのための現実的なバイノーラル音響データセット
AudioWorldSim: Realistic Binaural Audio Datasets For World Models
MetaのSoundSpaces 2.0を拡張し、ランダムなエージェント移動に基づく現実的なバイノーラル音響データセットを自動生成するオープンソースプラットフォームを提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Luis Vitor Zerkowski, Luiz Velho
分類: cs.SD, cs.LG
原文アブストラクト
This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. Built as a custom extension of Meta's SoundSpaces 2.0 platform, AudioWorldSim leverages their comprehensive acoustics framework, but focuses on the automatic rollout of random agent navigations, as well as implements crucial fixes to how continuous sound is composed. AudioWorldSim is made publicly available to the research community at https://github.com/Luizerko/AudioWorldSim to facilitate reproducibility.