日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.24749

D-JEPA: 意思決定に整合した潜在世界モデル

D-JEPA: A Decision-Aligned Latent World Model

シェア:XThreadsFacebookLINEはてブBluesky

実行結果から候補未来間の意思決定に関わる関係を学習し、潜在距離による計画を改善する世界モデルを提案。操作・自動運転・実機タスクで行動選択の成功率が向上。

詳しい要約

1. どんなもの?

- 潜在世界モデル(latent world model)の意思決定整合版 - 行動の結果を予測するが、予測精度だけでは実行成功を保証しない問題に対処 - D-JEPA は実行結果から候補未来間の意思決定関連関係を学習 - 目標相対の予測特徴と順序証拠を共同推論する有界・置換同変演算子 - 事前学習済み予測幾何を行動選択が重要な領域で精緻化 - JEPA 互換の未来表現で潜在距離計画を可能に - 潜在制御、操作、物理ロボット、自動運転で評価

2. 先行研究と比べてどこがすごい?

- 従来の潜在世界モデルは予測精度を重視 - しかし予測精度が高くても潜在距離が実行成功を反映しない「意思決定局所予測ギャップ」を指摘 - 目標に近いと予測された候補が、利用可能な代替より悪い結果を生む可能性 - D-JEPA は実行結果から意思決定関連関係を学習し、このギャップに対処 - 制限付き予測器適応と共有順序インターフェースで相補的予測幾何に整合を拡張 - 結果として PushT で 87.89% 成功、RoboTwin で平均 15.04 ポイント向上、物理ロボットで 17 ポイント向上

3. 技術・手法の肝は?

- 有界・置換同変演算子(bounded, permutation-equivariant operator) - 目標相対の予測特徴と順序証拠(ordinal evidence)を共同推論 - 事前学習済み予測幾何を行動選択が重要な領域で精緻化 - 制限付き予測器適応(restricted predictor adaptation) - 共有順序インターフェース(shared ordinal interface) - 学習した意思決定構造を JEPA 互換の未来表現で実現 - ネイティブな潜在距離計画(native latent-distance planning)で展開可能

4. どうやって有効だと検証した?

- 潜在制御、操作、事前学習済み行動生成モデル、物理ロボット、自動運転で評価 - PushT で 87.89% の成功率 - RoboTwin で平均 15.04 ポイントの向上 - 物理ロボットタスクで 17 ポイントの向上 - これらの結果が予測世界モデリングと効果的制御の橋渡しを確立

5. 議論はある?

- 意思決定関連の関係構造が予測世界モデリングと効果的制御の直接的な橋渡しになることを示す - 予測精度だけでは不十分である「意思決定局所予測ギャップ」の存在を強調 - 制限付き予測器適応と共有順序インターフェースの有効性 - 詳細な限界や失敗事例、計算コスト、スケーラビリティについては要旨からは不明

6. 次に読むべき論文は?

- JEPA(Joint Embedding Predictive Architecture) - 潜在世界モデル(latent world model) - モデルベース強化学習(model-based reinforcement learning) - 潜在距離計画(latent-distance planning) - 順序回帰(ordinal regression) - 置換同変ネットワーク(permutation-equivariant networks) - PushT、RoboTwin などのベンチマーク

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shuaijun Liu, Chengyu Wu, Qifu Wen, Feiyang You, Chenglong Zhang, Shuyang Hao, Xi Lin, Ningxin Su

分類: cs.RO, cs.LG

原文アブストラクト

Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.

関連論文

PR本紙発行元 EmplifAI