日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.03587

AVL-JEPA:世界モデルにおける因果ダイナミクス情報の崩壊を防ぐ

AVL-JEPA: Preventing Causal Dynamics Information Collapse In Joint Embedding Predictive Architecture World Models

シェア:XThreadsFacebookLINEはてブBluesky

JEPA世界モデルが視覚情報を保ちつつ行動の物理的結果を失う「因果ダイナミクス情報崩壊」を防ぐAVLを提案し、視覚摂動下でのロボット制御成功率を改善した。

詳しい要約

1. どんなもの?

JEPA(Joint Embedding Predictive Architecture)によるworld modelにおいて、観測を再構成せずに将来の潜在表現を予測する際に生じる『causal dynamics information collapse』(行動の物理的帰結に関する情報が失われる失敗モード)を防ぐ手法AVL(action-grounded vision-invariance latent)を提案した研究。 - JEPAは高次元視覚情報を保持しつつ、行動の因果的帰結情報を捨ててしまう問題がある。 - AVLはこの崩壊を防ぎ、計画に因果的に関連する動的情報を保持することを目指す。 - 4つのロボティック制御タスク(TwoRoom, PushT, OGBench Cube, Reacher)で検証。

2. 先行研究と比べてどこがすごい?

従来のJEPAは視覚情報を保持しつつも行動の物理的帰結情報を失う『causal dynamics information collapse』が生じる。 - AVLはこの崩壊を明示的に防ぎ、視覚摂動下でも成功率を大幅に改善しつつ、クリーン環境での性能を維持する点が優れている。 - 先行研究との具体的な比較数値は要旨からは不明。

3. 技術・手法の肝は?

AVLは2つの要素からなる。 - 実行されたactionを補助的なdynamics anchorとして用い、モデルに動的情報を保持させる。 - vision-invariance pathwayにより、摂動された潜在予測とクリーンな潜在予測を整合させ、動的情報を捨てずに因果的動的情報を完全に理解・利用するよう強制する。 - これによりcausal dynamics information collapseを防ぐ。

4. どうやって有効だと検証した?

4つのロボティック制御タスク(TwoRoom, PushT, OGBench Cube, Reacher)でAVLを検証。 - 視覚摂動下での成功率が大幅に向上し、クリーン環境での性能も維持されることを示した。 - さらにphysical consequence alignment、clean-noisy dynamics consistency、targeted transition subspace erasureの因果効果を評価。 - これらの結果から、計画に因果的に関連する動的情報がAVL下で崩壊から保持されることが示唆された。

5. 議論はある?

要旨からは不明。 - ただし、AVLが計画に因果的に関連する動的情報を崩壊から防ぐことを示す結果が得られている。 - 限界や今後の課題についての議論は要旨には記載されていない。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。 - 関連手法としてJEPA(Joint Embedding Predictive Architecture)が挙げられる。 - 同分野の定番としてworld model、model-based reinforcement learning、representation learningの論文を読むとよい。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yikang Qiao, Ling Zhang, Ziying Song, Duan Huang

分類: cs.RO

原文アブストラクト

Joint embedding predictive architectures (JEPAs) predict future latent representations without reconstructing observations, enabling world models to focus on high-level semantic dynamics. However, a JEPA can preserve high dimensional visual information while discarding information about the physical consequences of actions. We call this failure mode causal dynamics information collapse and propose action-grounded vision-invariance latent (AVL) to prevent this collapse. We first use the executed action as an auxiliary dynamics anchor that encourages the model to preserve dynamics information, and then use a vision-invariance pathway which aligns perturbed and clean latent predictions without discarding dynamics information, forcing the model to fully understand and utilize causal dynamics information. We validate AVL on four robotic control tasks (TwoRoom, PushT, OGBench Cube, and Reacher), showing that it substantially improves success rates under visual perturbations while preserving clean-environment performance. We further evaluate physical consequence alignment, clean-noisy dynamics consistency, and the causal effect of targeted transition subspace erasure. Collectively, these results indicate that dynamic information causally relevant to planning is preserved from collapse under AVL.

関連論文

PR本紙発行元 EmplifAI