日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
表現学習arXiv:2610.12349

SplitJEPA: 再構成なしで不変・変動潜在世界を学習する

SplitJEPA: Learning Invariant and Variant Latent Worlds without Reconstruction

シェア:XThreadsFacebookLINEはてブBluesky

再構成を行わずに、予測表現空間で潜在状態を不変部分と変動部分に分離して学習するJEPAを提案し、理論的な識別可能性を示した。

詳しい要約

1. どんなもの?

- 動的システムの潜在状態を、観測間で共有される不変因子と変化する変動因子に分解する手法。 - 再構成を行わず、表現空間で直接的に不変・変動部分を学習する JEPA の一種。 - ロボットのキューブ操作など、視点変化や照明変化に頑健な表現獲得を目指す。

2. 先行研究と比べてどこがすごい?

- 既存の分解手法は再構成を通じて潜在変数を得るため、観測全体を説明する必要があった。 - 既存の JEPA は再構成なしで潜在状態をモデル化するが、不変・変動構造の回復は達成されていない。 - SplitJEPA は再構成なしで不変・変動部分空間を同定できることを理論的に保証する点が新しい。

3. 技術・手法の肝は?

- JEPA の枠組みで、潜在状態とその不変・変動組織を同時に表現空間で学習。 - 定常ガウス予測ダイナミクスとフルランク変動条件の下で、不変・変動部分空間を独立なブロック単位の等長変換まで同定可能と証明。 - 観測デコーダを必要としないため、再構成なしの潜在回復を不変・変動ブロック同定に拡張。

4. どうやって有効だと検証した?

- 合成非線形システムとロボットマニピュレーションタスクで実験。 - 理論結果を支持し、ロバスト性と効率性の両面で実用的価値を示す。

5. 議論はある?

- 理論的保証は定常ガウス予測ダイナミクスとフルランク変動条件に依存。 - これらの仮定が現実の複雑なシステムでどの程度成立するかは要旨からは不明。 - 再構成なしで不変・変動を分離できる利点が、どのようなタスクで特に有効かは要旨からは不明。

6. 次に読むべき論文は?

- JEPA (Joint Embedding Predictive Architecture) に関する元論文。 - 再構成に基づく不変・変動分解の先行研究(例:β-VAE などの生成モデル)。 - ロボットマニピュレーションにおける表現学習の関連研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ruijin Hua, Zichuan Liu, Zhuokai Zhao, Yujia Zheng

分類: cs.LG

原文アブストラクト

Understanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that vary between them. For example, a robot pushing a cube to a goal should take the same action when the camera shifts or the lights dim, since nothing in the scene has moved. Existing approaches to this decomposition commonly obtain it through reconstruction, so the latent variables must first explain the entire observational world before their organization can be trusted. Joint embedding predictive architectures (JEPAs) model the latent state directly and never reconstruct, yet no existing result recovers the invariant and variant parts of the state they learn. How to learn the invariant-variant structure of the latent world without paying for its reconstruction therefore remains open. To close this gap, we introduce SplitJEPA, a JEPA that jointly recovers the latent state and its invariant and variant organization directly in representation space, without any reconstruction. We prove that, under stationary Gaussian predictive dynamics and a full-rank variation condition, SplitJEPA identifies the invariant and variant subspaces up to independent block-wise isometries, without introducing an observation decoder. Since the guarantee needs no decoder, the result extends reconstruction-free latent recovery to invariant-variant block identification. Experiments on synthetic nonlinear systems and robotic manipulation tasks support the theoretical results and show their practical value for both robustness and efficiency.

関連論文

PR本紙発行元 EmplifAI