合成状態遷移による動画インタラクション生成のブートストラップ
Bootstrapping Video Interaction Generation with Synthetic State Transitions
画像編集モデルで相互作用の開始・終了状態画像を作り、State-Guided Samplingで滑らかな動画を生成する合成データセットを構築し、ファインチューニングで物理的に妥当なインタラクション生成を改善した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim
分類: cs.CV
原文アブストラクト
While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit `start' and `end' state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.