H-JEPA: 視覚計画のための階層的世界モデルのエンドツーエンド学習
H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
各階層が独自の潜在空間でより遠い未来を予測するアクション条件付きJEPAの階層をエンドツーエンドで学習し、トップダウン計画により長期視覚計画の成功率と計算効率を改善する手法を提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Wancong Zhang, Basile Terver, Michael Rabbat, Yann LeCun, Randall Balestriero
分類: cs.LG, cs.RO
原文アブストラクト
Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down: the top level optimizes progress toward the goal, and each level's predictions become subgoals for the planner below it. When factors in the data evolve at separated timescales, higher levels discard fast, unpredictable detail and retain slower task-relevant state. Across four simulated navigation and manipulation environments, hierarchical planning improves over a flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% using less planner compute. Ablations attribute these gains to both temporal decomposition and higher-level goal representations. With inverse-dynamics supervision, the approach extends to diverse real-robot videos from DROID, where hierarchy improves offline planning fidelity at lower planner compute.