日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.06805

H-JEPA: 視覚計画のための階層的世界モデルのエンドツーエンド学習

H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning

シェア:XThreadsFacebookLINEはてブBluesky

各階層が独自の潜在空間でより遠い未来を予測するアクション条件付きJEPAの階層をエンドツーエンドで学習し、トップダウン計画により長期視覚計画の成功率と計算効率を改善する手法を提案。

詳しい要約

1. どんなもの?

- 長期horizonのvisual planningのための階層的world model「H-JEPA」を提案。 - action-conditioned JEPAを階層的にend-to-end学習。 - 各レベルは独自のlatent spaceでより遠い未来を予測。 - planningはtop-down:上位レベルがgoalへの進捗を最適化し、その予測が下位レベルのsubgoalとなる。 - データの時間スケールが分離しているとき、上位レベルは速く予測不能な詳細を捨て、遅いtask-relevant stateを保持。

2. 先行研究と比べてどこがすごい?

- 既存のtask-agnostic JEPA world modelは単一のtimescaleで予測・planningするか、複数のhorizonを一つのshared latent spaceで扱う。 - H-JEPAは階層ごとに異なるlatent spaceを持ち、各レベルがより遠い未来を予測。 - 4つのsimulated navigation/manipulation環境でflat JEPAより階層的planningが改善。 - Visual AntMazeでは3レベル階層がsuccessを18%から73%に向上させ、planner computeも削減。

3. 技術・手法の肝は?

- action-conditioned JEPAを階層的にend-to-endで学習するrecipe。 - 各レベルは自身のlearned latent spaceでより遠い未来を予測。 - planningはtop-down:最上位がgoalへの進捗を最適化し、各レベルの予測が下位plannerのsubgoalになる。 - 時間スケール分離時、上位レベルは速く予測不能な詳細を捨て、遅いtask-relevant stateを保持。 - inverse-dynamics supervisionにより、DROIDの多様なreal-robot動画に拡張可能。

4. どうやって有効だと検証した?

- 4つのsimulated navigation/manipulation環境で評価。 - Visual AntMazeで3レベル階層がsuccessを18%から73%に向上、planner compute削減。 - Ablationsでtemporal decompositionとhigher-level goal representationsの両方がgainに寄与すると特定。 - inverse-dynamics supervisionによりDROIDのreal-robot動画でoffline planning fidelityが向上し、planner computeも低減。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- JEPA (Joint Embedding Predictive Architecture) 関連のworld model研究。 - flat JEPA。 - DROID datasetを用いたreal-robot video研究。 - Visual AntMaze環境。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wancong Zhang, Basile Terver, Michael Rabbat, Yann LeCun, Randall Balestriero

分類: cs.LG, cs.RO

原文アブストラクト

Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down: the top level optimizes progress toward the goal, and each level's predictions become subgoals for the planner below it. When factors in the data evolve at separated timescales, higher levels discard fast, unpredictable detail and retain slower task-relevant state. Across four simulated navigation and manipulation environments, hierarchical planning improves over a flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% using less planner compute. Ablations attribute these gains to both temporal decomposition and higher-level goal representations. With inverse-dynamics supervision, the approach extends to diverse real-robot videos from DROID, where hierarchy improves offline planning fidelity at lower planner compute.

関連論文

PR本紙発行元 EmplifAI