LeWAM:拡散誘導MPCを用いたJEPA世界行動モデル
LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC
JEPA潜在空間上で順・逆・逆動力学と政策を双方向トランスフォーマで予測する世界行動モデルを提案し、政策ヘッドのノイズ空間でMPC計画を行うことで閉ループ性能を向上させた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Shashank Hegde, Alexander Popov, Elie Aljalbout, Nikolai Smolyanskiy
分類: cs.RO, cs.AI, cs.CV
原文アブストラクト
World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at matched size, while also providing a world model. 3) Planning: Sampling raw actions when planning with WAMs lets MPC exploit dynamics-model inaccuracies; planning in the noise space of the policy head instead improves the closed-loop performance of these WAMs.