日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画編集arXiv:2610.08779

ALIVE: 最初のフレームで誘導される動画編集のためのインタラクション整合型オブジェクト挿入

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

シェア:XThreadsFacebookLINEはてブBluesky

編集済みの最初のフレームと追加オブジェクトを指定する指示から、挿入オブジェクトが拾われる・操作されるなどのインタラクションに参加できる動画編集フレームワークを提案。

詳しい要約

1. どんなもの?

- ビデオ編集において、挿入したオブジェクトがビデオ内のアクションに参加する(例:拾われる、操作される)ことを可能にするフレームワーク「ALIVE」を提案。 - 編集された最初のフレームと、追加するオブジェクトのみを指定する指示を入力とする。 - 挿入オブジェクトがソースビデオの内容と一貫した相互作用を行うようにする。

2. 先行研究と比べてどこがすごい?

- 既存のビデオエディタはオブジェクトを挿入できるが、相互作用への参加が困難であった。 - ALIVEは、挿入オブジェクトを「生きている」ようにし、周囲のアクションを保持しつつ協調的なオブジェクト動作を実現。 - 2つのベンチマークで、最強の評価ベースラインと比較してOverallスコアをそれぞれ43.9%と4.4%改善。

3. 技術・手法の肝は?

- 35,800の編集ペアをキュレーション。3Dレンダリング、モデル生成、実世界ビデオとROSEの一般編集ペアを組み合わせ。 - 各ペアは対象オブジェクトの有無が異なり、周囲のアクションを保持。 - 視覚言語モデル(VLM)を訓練し、同じ入力から相互作用ガイダンスを予測。

4. どうやって有効だと検証した?

- ALIVE-interactionベンチマークを導入し、相互作用の忠実性、ソース保持、視覚的一貫性を統合VLMベースのプロトコルで評価。 - 一般的なビデオオブジェクト挿入ベンチマークでも評価。 - VLMガイダンスなしで、2つのベンチマークでOverallスコアが最強ベースラインより43.9%と4.4%向上。 - VLM予測ガイダンスにより、追加のユーザー入力なしでALIVE-interactionスコアが0.95ポイント向上。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- ROSE(一般編集ペアの提供元) - 一般的なビデオオブジェクト挿入ベンチマーク(具体的名称は要旨に記載なし)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou

分類: cs.CV

原文アブストラクト

Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects "alive" through coherent interactions with the source video's contents, using an edited first frame and an instruction naming only the added object. We curate 35,800 editing pairs combining 3D-rendered, model-generated, and real-world videos with general editing pairs from ROSE. Each pair differs in the target object's presence while preserving the surrounding action, teaching editors coordinated object behavior and source preservation. We further train a vision-language model (VLM) to predict interaction guidance from the same inputs. We introduce the ALIVE-interaction benchmark to assess interaction fidelity, source preservation, and visual coherence using a unified VLM-based protocol, and evaluate on the general video object insertion benchmark. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% and 4.4% on the two benchmarks, respectively. VLM-predicted guidance further improves the ALIVE-interaction score by 0.95 points without additional user inputs.

関連論文

PR本紙発行元 EmplifAI