日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
video-to-audioarXiv:2610.08760

WorldSonus: ワールドモデルに音を届ける

WorldSonus: Bringing Sound to Worlds

シェア:XThreadsFacebookLINEはてブBluesky

ワールドモデルが生成する映像に、リアルタイムで空間的なステレオ音声を合成する対話型のvideo-to-audioフレームワークを提案した論文。

詳しい要約

1. どんなもの?

- WorldSonusは、world modelsの生成環境に音を付与するinteractive video-to-audioフレームワーク。 - リアルタイム空間音響合成を目的とし、streaming causal autoregressive diffusion architectureを採用。 - 音声チャンクを低RTF 0.41で合成し、mid-streamの音指示に応答可能。 - ステレオ音響をシーン幾何とカメラ運動に整合させる。

2. 先行研究と比べてどこがすごい?

- 従来のworld modelsは視覚合成が中心で、生成環境はほぼ無音だった。 - 既存のvideo-to-audioモデルは双方向型が主流で、リアルタイム性・対話制御・空間整合の同時達成が困難。 - WorldSonusはworld models向けに設計しつつ、open-domain video-to-audioベンチマークでも双方向SOTAモデルに匹敵または上回る。

3. 技術・手法の肝は?

- streaming causal autoregressive diffusion architectureで音声チャンクを逐次合成。 - audio-centric captioning pipelineとchunk-indexed prompt schedulingにより、生成中の音イベントを動的に操作。 - 多様なstereoおよびambisonicデータからキュレーションした高品質ステレオ教師信号を利用し空間整合を実現。

4. どうやって有効だと検証した?

- 広範な実験を実施し、world models向けでありながらopen-domain video-to-audioベンチマークで汎化を確認。 - 音響品質と空間整合の両面で、双方向SOTAモデルと同等以上に性能を発揮。 - リアルタイム性はRTF 0.41で示される。

5. 議論はある?

- 要旨からは、限界や失敗事例、計算コスト、データ依存性などの議論は不明。 - リアルタイム性・対話制御・空間整合の三課題を同時に扱う点が主な主張。

6. 次に読むべき論文は?

- 要旨で参照・比較されている具体的な先行研究名は不明。 - 関連手法として、world models、video-to-audio synthesis、autoregressive diffusion models、stereo/ambisonic audio processingの定番文献を次に読むべき。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang, Haoyuan Huang, Yaguang Wu, Xiangjun Huang, Ziyang Ma, Weijia Chen, Hongyu Liu, Zeyue Tian, Qifeng Chen

分類: cs.SD, cs.AI, cs.CV, eess.AS

原文アブストラクト

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/

PR本紙発行元 EmplifAI