日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.22403

LD4WAM: 人間のビデオから潜在ダイナミクスを学習するワールドアクションモデル

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

人間のビデオからロボットの行動に直接使える潜在ダイナミクスを学習し、将来のビデオ生成と行動条件付けを組み合わせたワールドアクションモデルを提案。5,000時間以上のデータで事前学習し、シミュレーションと実機で高い性能を達成。

詳しい要約

1. どんなもの?

LD4WAMは、人間のビデオから学習した潜在ダイナミクスを利用して、ロボットの行動生成を改善するWorld Action Model (WAM) の手法である。人間のビデオは多様性があり収集コストが低いが、ピクセルレベルの未来予測は直接行動に利用できない。そこで、embodimentに依存しない運動整合的な潜在ダイナミクス表現を導入し、ビデオの事前知識と低レベルの行動を橋渡しする。LD4WAMは、Latent Dynamics ModelとWorld Dynamics Action Modelを組み合わせ、5,000時間以上の人間とロボットのデータで事前学習され、シミュレーションと実ロボットで有効性を示す。

2. 先行研究と比べてどこがすごい?

従来のWAMは人間のビデオからピクセルレベルの未来フレームを予測するだけで、行動に直接利用できないダイナミクスしか得られなかった。また、motion retargetingは直接行動を復元できるが、embodiment間の視覚的ギャップが大きい。LD4WAMは、運動整合的な潜在ダイナミクスを導入することで、ビデオの事前知識と低レベル行動のギャップを埋め、embodimentに依存しない表現を実現した点が新しい。

3. 技術・手法の肝は?

手法の核は、意味的再構成と実運動整合で訓練されたLatent Dynamics Modelと、mixture-of-transformers (MoT) で構成されたWorld Dynamics Action Modelの組み合わせである。MoTは未来ビデオ生成を保持しつつ、学習可能なクエリを用いて生成された未来から潜在ダイナミクスを抽出し、行動条件付けに利用する。これにより、人間のビデオから得た潜在ダイナミクスをロボットの行動生成に活用する。

4. どうやって有効だと検証した?

5,000時間以上の人間とロボットのデータを含む統一データセットで事前学習し、RoboTwinシミュレーションと、グリッパーと器用な手を備えた実ロボットで評価した。未見の物体や背景に対する汎化性能も検証し、強い性能を示した。

5. 議論はある?

要旨からは、潜在ダイナミクスの表現がどの程度embodimentに依存しないか、また、実ロボットでの具体的な成功率やシミュレーションとの差などの詳細は不明。また、データセットの構成や事前学習の詳細、計算コストなども要旨からは不明である。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、World Action Models (WAMs)、motion retargeting、mixture-of-transformers (MoT) が挙げられる。また、RoboTwinシミュレーションや、人間ビデオからのロボット学習に関する既存研究(例:Learning from Human Videos)も関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhenhao Shen, Jiaqi Liang, Jasper Lu, Feng Jiang, Yuran Wang, Chuanbo Wei, Jiayi Liu, Jianchun Yang, Qize Yu, Jiadi You, Ce Hao, Guanqi He, Chen Xie, Ruihai Wu

分類: cs.RO

原文アブストラクト

Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. However, most WAMs learn from such video only by predicting pixel-level future frames, giving dynamics that are not directly actionable, whereas motion retargeting recovers directly actionable actions but leaves a large visual gap across embodiments. We therefore propose motion-aligned latent dynamics as an embodiment-agnostic representation to bridge video priors and low-level actions. We further present LD4WAM, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated futures for action conditioning. Pretrained on our curated unified dataset of over 5{,}000 hours of human and robot data, LD4WAM performs strongly in RoboTwin simulation and on real robots equipped with both grippers and dexterous hands, while generalizing well to unseen objects and backgrounds.

関連論文