日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.06598

SimForcing: シミュレーションの運動事前分布を実世界ロボット世界モデルへ蒸留

SimForcing: Distilling Simulation Motion Priors into Real-Domain Robot World Models

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーションの運動知識を潜在空間蒸留で実動画生成モデルに転移し、シミュレーション予測を条件として活用することで、実世界ロボットの高品質な行動条件付き世界モデルを実現した。

詳しい要約

1. どんなもの?

- シミュレーションを活用したロボット世界モデル学習フレームワーク - 実世界のロボット動画から行動条件付き世界モデルを学習 - シミュレーションの運動知識を蒸留し、外観差を緩和 - 推論時に追加の世界モデル不要で、シミュレーション条件と実世界動画を生成 - BridgeやInternData-A1で評価し、下流のVLAモデル初期化にも有効

2. 先行研究と比べてどこがすごい?

- 従来の実世界動画のみからの学習と比べ、シミュレーションの構造化された運動監督を利用 - 外観差による直接転移の困難をlatent-motion distillationで克服 - シミュレーション予測の不正確さに過度に依存しないmulti-block simulation conditioningとcondition dropoutを導入 - 外部のembodied pretrainingなしでBridge上で最高のPSNR, SSIM, LPIPS, FVDを達成

3. 技術・手法の肝は?

- シミュレーション教師からlatent-motion distillationで運動知識を転移 - 潜在空間の時間変化を整合させ、外観差の影響を軽減 - multi-block simulation conditioningとcondition dropoutでシミュレーション軌道を活用 - simulation-conditioning classifier-free guidanceで内部化した運動知識とシミュレーション潜在による予測を統合 - 学生モデルはシミュレーション条件と実世界動画を共同で生成

4. どうやって有効だと検証した?

- BridgeデータセットでPSNR, SSIM, LPIPS, FVDを比較し、外部embodied pretrainingなしで最高性能 - InternData-A1でも評価し、ロボットデータセット間の適用性を確認 - 学習した世界モデルでvision-language-actionモデルを初期化し、LIBERO成功率が向上

5. 議論はある?

- シミュレーション予測の不正確さが実動画生成を誤導する可能性を指摘 - 外観差の影響をlatent-motion distillationで軽減するが、完全な解決かは要旨からは不明 - 推論時に追加の世界モデルが不要である点を強調 - 下流ポリシー学習への有用性を示唆するが、詳細な議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 明示的な論文名はなし - 関連手法: action-conditioned robot world models, latent-motion distillation, classifier-free guidance, vision-language-action models - 同分野の定番: Dreamer, RoboDreamer, UniSim, World Models

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiaodong Wang, Tianle Li, Chuanxin Song, Junliang Xie, Zhanmi Zhong, Suiying Wu, Peixi Peng

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation both as a source of transferable motion knowledge and as a controllable reference for prediction. First, we transfer motion knowledge from a simulation teacher through latent-motion distillation, aligning temporal changes in latent space to internalize motion priors while mitigating the influence of appearance differences. Second, we introduce multi-block simulation conditioning with condition dropout to exploit predicted simulation trajectories without relying excessively on their accuracy. Our simulation-conditioning classifier-free guidance scheme unifies these two ideas by balancing predictions based on internalized motion knowledge with those additionally guided by simulation latents. The jointly trained student generates both simulation conditions and real-domain videos, requiring no additional world model at inference. On Bridge, SimForcing achieves the best PSNR, SSIM, LPIPS, and FVD among the compared methods without external embodied pretraining. Evaluation on InternData-A1 further supports its applicability across robot datasets. Moreover, using our trained world model to initialize a vision-language-action model improves LIBERO success, suggesting its utility for downstream policy learning. \url{https://github.com/Wang-Xiaodong1899/SimForcing}

関連論文

PR本紙発行元 EmplifAI