SimForcing: シミュレーションの運動事前分布を実世界ロボット世界モデルへ蒸留
SimForcing: Distilling Simulation Motion Priors into Real-Domain Robot World Models
シミュレーションの運動知識を潜在空間蒸留で実動画生成モデルに転移し、シミュレーション予測を条件として活用することで、実世界ロボットの高品質な行動条件付き世界モデルを実現した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Xiaodong Wang, Tianle Li, Chuanxin Song, Junliang Xie, Zhanmi Zhong, Suiying Wu, Peixi Peng
分類: cs.RO, cs.AI, cs.CV
原文アブストラクト
Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation both as a source of transferable motion knowledge and as a controllable reference for prediction. First, we transfer motion knowledge from a simulation teacher through latent-motion distillation, aligning temporal changes in latent space to internalize motion priors while mitigating the influence of appearance differences. Second, we introduce multi-block simulation conditioning with condition dropout to exploit predicted simulation trajectories without relying excessively on their accuracy. Our simulation-conditioning classifier-free guidance scheme unifies these two ideas by balancing predictions based on internalized motion knowledge with those additionally guided by simulation latents. The jointly trained student generates both simulation conditions and real-domain videos, requiring no additional world model at inference. On Bridge, SimForcing achieves the best PSNR, SSIM, LPIPS, and FVD among the compared methods without external embodied pretraining. Evaluation on InternData-A1 further supports its applicability across robot datasets. Moreover, using our trained world model to initialize a vision-language-action model improves LIBERO success, suggesting its utility for downstream policy learning. \url{https://github.com/Wang-Xiaodong1899/SimForcing}
関連論文
- 低コストロボットナビゲーションにおける効率的なSim-to-Real転移のためのデュアル変分オートエンコーダsim2real
- 人間の動画を物理的に整合したロボット操作データに変換sim2real
- 視覚ベースUAV着陸における二項結果を伴うDNN再学習のためのベイズデータ拡張sim2real
- ArtifactArena:物理世界で何を作れるかでモデルを評価するsim2real
- AffordCraft: 単一画像からタスク対応シミュレーション資産をスケーラブルに構築sim2real
- サンプリングベース外乱オブザーバによるSim-to-Realギャップの克服:解析モデルから学習型世界モデルまでsim2real