日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画生成arXiv:2609.38154

LongLive-Plug: 動画生成のための一度きりの蒸留フレームワーク

LongLive-Plug: Once-for-All Distillation for Video Generation

シェア:XThreadsFacebookLINEはてブBluesky

ベースモデルにLoRAとして再利用可能な能力を学習させ、下流モデルに訓練不要でプラグアンドプレイ展開できる蒸留手法を提案。

詳しい要約

1. どんなもの?

ビデオ拡散モデル向けの一度きりの蒸留フレームワーク LongLive-Plug を提案する。 - ベースモデル上で再利用可能な能力を LoRA として学習し、互換性のある下流モデルへ training-free かつ plug-and-play で適用する。 - 対象能力は single-pass classifier-free guidance、few-step sampling、autoregressive generation の long-context error correction。 - 下流モデルが conditioning branch 追加や output channel 拡張をしても adapter は再利用可能。 - 3 backbone families、8 task categories、54 下流モデルで training-free 展開を検証。

2. 先行研究と比べてどこがすごい?

従来は下流タスクごとに蒸留段階を繰り返す必要があった。 - LongLive-Plug は backbone family ごとに一度蒸留すれば、対象ごとの再学習なしで再利用できる。 - 下流モデルの構造変更(conditioning branch 追加、output channel 拡張)後も adapter が再利用可能。 - 固定 guidance scale で学習しても、CFG LoRA の inference weight で text guidance を制御できる。 - few-step LoRA と組み合わせても few-step 生成と CFG 制御性を両立。

3. 技術・手法の肝は?

ベースモデルに LoRA として再利用可能な能力を蒸留する once-for-all フレームワーク。 - 能力ごとに専用 LoRA を用意:single-pass CFG、few-step sampling、long-context error correction。 - 下流モデルへは training-free で plug-and-play 適用。 - CFG LoRA は固定 guidance scale で学習されるが、推論時の重みで text guidance を制御。 - few-step LoRA と併用しても few-step 生成と CFG 制御性を維持。 - 下流モデルの conditioning branch 追加や output channel 拡張に対しても adapter が再利用可能。

4. どうやって有効だと検証した?

3 backbone families、8 task categories、54 下流モデルで training-free 展開を検証。 - タスクカテゴリには world modeling、robotics、editing、multimodal generation を含む。 - 各能力は backbone family ごとに一度蒸留し、対象ごとの再学習なしで再利用できることを確認。 - 追加の互換モデルにも対応し得ると述べる。

5. 議論はある?

要旨からは不明。 - 限界や失敗事例、計算コスト、LoRA の互換性条件などは明記されていない。 - 追加の互換モデルへの対応可能性は示唆されるが、詳細な議論は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている個別研究は明示されていない。 - 関連手法として video diffusion models、classifier-free guidance、LoRA、few-step sampling、autoregressive generation、long-context error correction が挙げられる。 - 同分野の定番として diffusion distillation、consistency models、adversarial distillation などが次に読む候補。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

分類: cs.CV

原文アブストラクト

Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.

関連論文

PR本紙発行元 EmplifAI