日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
音声・映像編集arXiv:2605.18467

指示に基づく音声・映像同時編集フレームワークInstructAV2AV

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

シェア:XThreadsFacebookLINEはてブBluesky

指示文に従って音声と映像を同時に編集する初のエンドツーエンドフレームワークを提案し、大規模データセットと新しい注意機構で高品質な編集を実現した。

著者: Haojie Zheng, Yixin Yang, Siqi Yang, Shuchen Weng, Boxin Shi

分類: cs.CV

原文アブストラクト

Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose InstructAV2AV, the first end-to-end framework for instruction-guided audio-video joint editing. We first develop a scalable data synthesis pipeline and construct InsAVE-80K, the first large-scale audio-video editing dataset with high-quality source-to-target pairs. With this data foundation, we adapt an audio-video generation backbone to leverage its robust priors. We concatenate the audio-video input with noisy latent codes to anchor the source context, propose the source-instruction gated attention to improve instruction following and content preservation, and introduce a two-stage training strategy to effectively transfer these pre-trained priors. Extensive experiments demonstrate that InstructAV2AV outperforms state-of-the-art methods across 11 metrics spanning three aspects on two evaluation sets, highlighting its potential for controllable content creation. Project page: https://hjzheng.net/projects/InstructAV2AV/.