動画タスクを時空間アナロジーで統合するViGeo
Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy
画像分野に限られていた視覚的アナロジーによる文脈内学習を動画領域に拡張し、時空間キャンバス補完によって多様な動画タスクを学習なしで統合・汎化するフレームワークViGeoを提案。
著者: Chia-Hsiang Kao, Belinda Zeng, Bharath Hariharan, Menglin Jia
分類: cs.CV, cs.AI
原文アブストラクト
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.