日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.10515

RoboJEPA: ロボット潜在世界モデルのスケーリング

RoboJEPA: Scaling Robotic Latent World Models

シェア:XThreadsFacebookLINEはてブBluesky

12種類のロボット実機データでJEPAベースの世界モデルを大規模学習し、計算量に対するスケーリング則を確立するとともに、ゼロショットで長期計画タスクを実機実行できることを示した。

詳しい要約

1. どんなもの?

- ロボットの潜在世界モデル「RoboJEPA」を提案。 - Joint Embedding Predictive Architecture (JEPA) に基づく。 - 12種類のembodimentを含む大規模データセットで学習。 - 8Bパラメータの最大規模のJEPA predictor。 - 実ロボットデータで学習したマルチembodiment世界モデルのscaling lawを初めて確立。

2. 先行研究と比べてどこがすごい?

- 従来、潜在世界モデルの能力がモデルサイズ・データ・計算量でどうスケールするか不明だった。 - RoboJEPAはimagination errorがcomputeに対して二次のpower lawに従うことを示す。 - 下流のplanning性能もcomputeとともに予測可能に向上。 - imagination errorが実ロボット評価の信頼できる代理指標となることを示す。 - 実ロボットデータで学習したマルチembodiment世界モデルのscaling lawは初。

3. 技術・手法の肝は?

- JEPAベースの世界モデル。 - 12 embodimentの大規模データセットで学習。 - 潜在rolloutの誤差(imagination error)を指標化。 - computeに対するscaling lawをフィッティング。 - 単一のgoal imageに向けてplanningするzero-shotエージェントとして展開。

4. どうやって有効だと検証した?

- imagination errorがcomputeの二次のpower lawに従うことを示す。 - 下流のplanning性能がcomputeとともに予測可能に向上することを確認。 - imagination errorと実ロボット性能の強い相関を検証。 - 実ハードウェアで長期的planningを要するタスクをzero-shotで解決。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、汎化性に関する議論は明記されていない。

6. 次に読むべき論文は?

- JEPA (Joint Embedding Predictive Architecture) 関連の論文。 - 世界モデルとplanningのscaling lawに関する研究。 - マルチembodimentロボット学習のデータセット論文。 - 具体的な参照論文は要旨に明記されていないため、同分野の定番としてJEPAや世界モデル、ロボット学習の代表的文献を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Artem Zholus, Nicolas Beltran-Velez, Jianhao Yuan, Sarath Chandar, Tushar Nagarajan, Daniel Severo, Koustuv Sinha, Michal Drozdzal, Adriana Romero Soriano, Jeannette Bohg, Nicolas Ballas, Mahmoud Assran

分類: cs.AI, cs.RO

原文アブストラクト

Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a world model based on the Joint Embedding Predictive Architecture (JEPA) and trained on a large-scale dataset spanning 12 robotic embodiments. We show that RoboJEPA's imagination error, the error of its latent rollouts, follows a second-order power law in compute, allowing us to predict model quality well beyond the scale at which the law is fit. We further show that downstream robotic planning performance improves predictably with compute, and that imagination error is strongly correlated with it, making it a reliable proxy for real-robot evaluation. Finally, we demonstrate that latent world models can be deployed zero-shot as robotic agents, planning toward a single goal image to solve tasks requiring long-horizon planning on real hardware. We release all model checkpoints together with our training and robot deployment code. To our knowledge, this is the first work to establish scaling laws for multi-embodiment robotic world models trained on real robot data, and RoboJEPA, at 8B parameters, is the largest JEPA predictor model trained to date.

関連論文

PR本紙発行元 EmplifAI