日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
動画生成/蒸留arXiv:2609.31349

DyMD: 分布マッチング蒸留による少数ステップ動画世界モデルでの相互作用ダイナミクス保持

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

シェア:XThreadsFacebookLINEはてブBluesky

動画拡散モデルの蒸留において、ロボットと物体の動きを保ちながら4ステップで生成できるDyMDを提案し、身体性動画ベンチマークでタスク遵守率を9.6ポイント改善した。

詳しい要約

1. どんなもの?

- 大規模video diffusion modelのmany-step samplingは対話的下流利用で高コスト。 - Distribution Matching Distillation (DMD)でfew-step video generationが可能だが、robot-object motionを抑制しvisual qualityを維持する問題。 - 提案DyMDは、teacher supervisionとcritic fittingを進化するstudentに適応させるDMD framework。 - 14B teacherをfour-step 1.3B studentに蒸留し、推論時auxiliary modules不要。 - embodied-video benchmarksでR-Bench task adherenceとPAI-Bench-G Domain scoreを改善。

2. 先行研究と比べてどこがすごい?

- Base DMDはfew-step video generationを可能にするが、robot-object motionを抑制。 - DyMDはteacher posteriorがmotion-deficient rolloutsに集中する問題と、stronger-motion rolloutsのfake-score fitting errorsを解決。 - 結果として、Base DMD比でR-Bench task adherenceが9.6 percentage points、PAI-Bench-G Domain scoreが5.1 points改善。 - 下流action planningのbackboneとして、WorldArenaタスクでmean success 34% (Base DMDは16%)。

3. 技術・手法の肝は?

- Temporal affinity-conditioned re-noise sampling: 各rolloutのinteraction fidelityに応じてtimestep distributionを適応。 - base scheduleとteacher priorを混合し、motion recoveryとappearance refinementをバランス。 - Dynamics-guided fake-score tracking: noise-conditioned predictorでnoise-relative fitting difficultyを推定。 - latent temporal dynamicsから予測し、critic lossでpredicted-hard rolloutsをupweight。 - 推論時auxiliary modules不要で、14B teacherからfour-step 1.3B studentへ蒸留。

4. どうやって有効だと検証した?

- embodied-video benchmarksで評価。 - R-Bench task adherence: Base DMD比+9.6 percentage points。 - PAI-Bench-G Domain score: Base DMD比+5.1 points。 - visual qualityは同等を維持。 - 下流action planningのbackboneとしてWorldArenaタスクで評価。 - mean success: 34% (Base DMDは16%)。

5. 議論はある?

- DMDのteacherとfake-score signalsを分析し、weak re-noisingがteacher posteriorをmotion-deficient rollouts近傍に集中させ、motion-restoring guidanceを制限。 - stronger-motion rolloutsはfake-score fitting errorsが大きくなり、generatorのinteraction dynamics学習を妨げる可能性。 - これらの問題に対処するため、teacher supervisionとcritic fittingを適応させるDyMDを提案。 - 要旨からは、他の議論や限界は不明。

6. 次に読むべき論文は?

- Distribution Matching Distillation (DMD) の原論文。 - 比較対象のBase DMD。 - 関連するvideo diffusion modelsのfew-step生成手法。 - embodied predictionやaction planningのためのvideo world models。 - 要旨で参照されているR-Bench、PAI-Bench-G、WorldArenaの各ベンチマーク論文。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu

分類: cs.CV, cs.AI

原文アブストラクト

Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.

PR本紙発行元 EmplifAI