日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.34375

LRC-JEPA:ダイナミクスと残差コンテキストを分離した効率的な世界モデル

LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models

シェア:XThreadsFacebookLINEはてブBluesky

予測に必要な動的状態と時間的に持続する視覚的文脈を別々の潜在表現に分離することで、軽量かつ高精度なJEPA型世界モデルを実現した。

詳しい要約

1. どんなもの?

LRC-JEPAは、JEPA系のworld modelにおける表現の絡み合い問題を解決する軽量なend-to-end world model。 - 情報を2つに分離: 予測latent z と learned-query residual-context embedding u。 - zのみをdynamics modelで伝播しplanningに使用。 - uは時間的に持続する情報を捉え、cross-attention reconstructionに使う。 - differentiable residual connectionでzに相補的な動的contentを保持させる。 - 4つのsimulated control環境と実世界Bridge-v2で評価。

2. 先行研究と比べてどこがすごい?

パラメータ数を揃えたJEPA baselineに対し平均planning成功率を9ポイント改善。 - 大幅に大きいpretrained modelと同等以上。 - 実世界Bridge-v2で5.5Mパラメータのactive encoderがDINO-WM(22.1M)やV-JEPA2(303.9M)のencoderを上回る。 - より高速なplanningも可能。 - 低次元表現で制御可能状態とnuisance appearanceの絡み合いを避ける点が先行研究と異なる。

3. 技術・手法の肝は?

情報をcompact predictive latent zとlearned-query residual-context embedding uにルーティング。 - zのみをdynamics modelで伝播しplanningに使用。 - uは時間的に持続する情報を捉え、cross-attention reconstructionに用いる。 - differentiable residual connectionによりzが相補的な動的contentを保持するよう促す。 - 明示的な仮定の下で、表現がsufficient, minimal, nuisance-invariant, disentangledであることを示す。

4. どうやって有効だと検証した?

4つのsimulated control環境でplanning成功率を評価。 - パラメータ一致のJEPA baselineと比較し平均9ポイント改善。 - 実世界Bridge-v2 setで5.5Mパラメータのactive encoderがDINO-WM(22.1M)とV-JEPA2(303.9M)を上回る。 - physical-state probes, reconstruction interventions, ablationsで表現のdisentanglementの有効性を確認。

5. 議論はある?

要旨からは不明。 - 明示的な仮定の下での理論的性質(sufficient, minimal, nuisance-invariant, disentangled)が示されているが、仮定の妥当性や限界についての議論は要旨からは不明。 - 実世界Bridge-v2での結果は示されているが、他の実世界タスクへの一般化や計算コストの詳細は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究: JEPA baseline, DINO-WM, V-JEPA2。 - 関連手法としてJEPA系world model, DINO-WM, V-JEPA2を挙げる。 - 同分野の定番としてmodel-based reinforcement learningやlatent-space planningの研究も参考になる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao

分類: cs.LG, cs.AI, cs.RO

原文アブストラクト

Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.

関連論文

PR本紙発行元 EmplifAI