日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
生成AI・インタラクティブアートarXiv:2609.27489

通過:AI生成音と再構成された時空間を巡る終わりなき旅

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

シェア:XThreadsFacebookLINEはてブBluesky

単一のモノレール車窓映像を時空間ボリュームとして再構成し、視聴者の存在に応じて非線形に再サンプリングすることで、終わりのない風景とAI生成音を生み出すインタラクティブな視聴覚インスタレーション。

詳しい要約

1. どんなもの?

本論文は、単一の連続したモノレール車窓録画を時空間ボリュームとして再構成し、終わりのない旅を生成するインタラクティブな音響視覚インスタレーション「Passing」を紹介する。映像を線形に再生するのではなく、その空間的・時間的構造を非線形な軌跡に沿ってリサンプリングし、奥行き・速度・時間的順序が不安定になる連続的な風景を生み出す。カメラベースの視聴者存在検出システムが視聴ゾーン内の視聴者の有無を推定し、その存在状態がレンダリングされた映像シーケンス間の遷移に影響を与える。得られた映像ストリームはリアルタイムのvideo-to-audio合成モデルSpecMaskFoleyに入力され、再構成された映像に同期したsoundscapeを生成する。

2. 先行研究と比べてどこがすごい?

要旨からは不明。従来のvideo-to-audio合成やインタラクティブインスタレーションとの具体的な比較は述べられていない。

3. 技術・手法の肝は?

技術の肝は、単一の連続したモノレール車窓録画を時空間ボリュームとして再構成し、空間的・時間的構造を非線形な軌跡に沿ってリサンプリングすることで、奥行き・速度・時間的順序が不安定な連続的風景を生成する点にある。さらに、カメラベースのviewer-presence detection systemが視聴ゾーン内の視聴者の有無を推定し、その存在状態がレンダリングされた映像シーケンス間の遷移に影響を与える。得られた映像ストリームはリアルタイムのvideo-to-audio合成モデルSpecMaskFoleyに入力され、同期したsoundscapeを生成する。SpecMaskFoleyは客観的に正しいsoundtrackを再構成するのではなく、speculative listenerとして機能し、従来の空間的・時間的前提が崩れた世界に対する可能な聴覚的解釈を提案する。

4. どうやって有効だと検証した?

要旨からは不明。有効性の検証方法については述べられていない。

5. 議論はある?

本作品は、時空間再構成の規則を定義するアーティスト、創発的な視覚フローを音として解釈するAIモデル、そして身体的存在が音響視覚的軌跡に影響を与える観客の間で創造的エージェンシーを分散させる。この構造を通じて、人間の意図、機械の推論、観客の解釈の間で著作者性と聴取がどのように交渉され得るかを探求する。

6. 次に読むべき論文は?

要旨で参照されているSpecMaskFoley。関連手法としてvideo-to-audio synthesis、audiovisual installation、spatiotemporal reconstructionが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Akira Takahashi, Chihiro Nagashima, Zhi Zhong, Shusuke Takahashi, Yuki Mitsufuji

分類: cs.SD, cs.AI

原文アブストラクト

This paper introduces Passing, an interactive audiovisual installation that generates an endless journey from a single continuous monorail-window recording by reconstructing it as a spatiotemporal volume. Rather than replaying the footage linearly, the work resamples its spatial and temporal structure along nonlinear trajectories, producing a continuously passing landscape whose depth, speed, and temporal order become unstable. A camera-based viewer-presence detection system estimates whether a viewer is present in the viewing zone and uses this presence state to influence transitions among rendered video sequences. The resulting video stream is fed into SpecMaskFoley, a real-time video-to-audio synthesis model that generates a synchronized soundscape for the reconfigured image. The model is not used to reconstruct an objectively correct soundtrack, but functions as a speculative listener, proposing a possible auditory interpretation of a world whose conventional spatial and temporal premises have been disrupted. Passing distributes creative agency across the artist, who defines the rules of spacetime reconstruction; the AI model, which interprets the emergent visual flow as sound; and the audience, whose embodied presence influences the audiovisual trajectory. Through this structure, the work investigates how authorship and listening may be negotiated among human intention, machine inference, and audience interpretation. Artwork page: https://ryufurusawa.com/passing

PR本紙発行元 EmplifAI