日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.36413

無限の未来から一つを実現:事前学習済み世界モデルをロボット行動へ変換

One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions

シェア:XThreadsFacebookLINEはてブBluesky

凍結した動画世界モデルの上に60Mパラメータの小型モデルを載せ、タスクに合う未来を選んで行動を読み出すことで、単一GPUで学習できリアルタイム制御を可能にした。

詳しい要約

1. どんなもの?

- ビデオ世界モデル(video world model)をロボット行動に変換する枠組み - 事前学習済み世界モデルは多数の未来を許容するが、ロボットはタスク条件に合う特定の未来を実行する必要 - この課題を actualization と定式化し、凍結した世界モデルの上にタスク条件付きの選択と実現を学習 - 実装は RoboActualizer で、60M パラメータの小型モデル - 2 つの軽量 DiT experts が flow matching で未来潜在と行動を同時予測 - 単一 GPU(32GB ピークメモリ)で学習可能、39ms の低遅延でリアルタイム制御を実現

2. 先行研究と比べてどこがすごい?

- 既存手法は大規模ロボットデータと計算資源で重い世界モデル backbone を fine-tune - 本研究は「高コストな部分は世界モデル事前学習で既に支払い済み」と主張 - 凍結した世界モデルを活用し、選択と読み出しのみを学習するため、学習可能パラメータが最大 100 倍少ない - 既存の WAMs や VLAs と比較して大幅に軽量 - 単一 GPU で学習可能で、計算コストを劇的に削減

3. 技術・手法の肝は?

- actualization をタスク条件付きの選択と実現の学習として定式化 - 凍結した世界モデル encoder の上に小型 actualizer を構築 - actualizer は 2 つの軽量 DiT experts から構成 - flow matching により未来潜在と行動を同時に予測 - タスク条件に基づき、世界モデルが示す多様な未来から実行すべき未来を選択し、行動に読み出す

4. どうやって有効だと検証した?

- シミュレーションベンチマーク LIBERO, LIBERO-Plus, RoboTwin 2.0 で評価 - 2 つの実世界プラットフォーム上の 5 タスクで評価 - 39ms の低遅延でリアルタイム制御が可能であることを確認 - 学習可能パラメータが最大 100 倍少ないにもかかわらず高い性能を達成

5. 議論はある?

- 世界モデル事前学習の表現空間が多様な未来を既に含むという仮定に基づく - 凍結した世界モデルで十分な選択・実現が可能であることを示唆 - 計算資源制約の緩和と実時間制御の両立を主張 - ただし、世界モデルの品質やタスク多様性への依存性、汎化限界については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: WAMs, VLAs - 関連手法: flow matching, DiT (Diffusion Transformer) - ベンチマーク: LIBERO, LIBERO-Plus, RoboTwin 2.0 - 同分野の定番: video world model, robot policy learning, fine-tuning of large-scale models

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bang Du, Yichen Xie, Shuqi Zhao, Yuxin Chen, Menglin Wu, Masayoshi Tomizuka

分類: cs.RO

原文アブストラクト

A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures. In this case, what remains is to select the future that accomplishes the task and to read out the actions that realize it. We formalize this task as actualization, which learns a task-conditioned selection and realization on top of a prior supplied by a frozen world model. This can be solved by a tiny actualizer model. We implement RoboActualizer with as few as 60M parameters on top of a frozen world model encoder. The actualizer is composed of two lightweight DiT experts that jointly predict future latents and actions by flow matching. The model can be trained entirely on a single GPU with 32 GB peak memory. With up to 100x fewer trainable parameters than existing WAMs and VLAs, RoboActualizer reaches great performance on simulation benchmarks including LIBERO, LIBERO-Plus, RoboTwin 2.0 and five tasks on two real-world platforms, with a low latency of 39 ms that allows real-time control.

関連論文

PR本紙発行元 EmplifAI