日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.07540

JEPA世界モデルにおける逆動力学による不安定モードの保持

Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models

シェア:XThreadsFacebookLINEはてブBluesky

JEPA世界モデルに逆動力学損失を追加することで、制御に重要な不安定モードを表現に保持し、視覚ベースの制御を可能にする手法を提案。

詳しい要約

1. どんなもの?

- ロボティクスにおける不安定モードを保持する表現学習の研究。 - JEPA(Joint-embedding predictive architecture)を用いたworld modelにおいて、次ステップ予測とanti-collapse正則化だけでは制御可能な不安定モードが潰れる問題を指摘。 - 逆動力学損失(action reconstruction objective)を追加することで、制御に重要な特徴を保持するcontrol-aware表現を学習。 - 線形系で理論解析し、非線形視覚制御タスク(CartPole, Walker2D, PointMaze)で有効性を実証。

2. 先行研究と比べてどこがすごい?

- 従来のJEPAベースのworld modelは次ステップ予測とanti-collapse正則化で表現を学習するが、制御可能な不安定モードが潰れる可能性を理論的に示した点が新しい。 - 逆動力学損失を追加することで、エンコーダが有限ホライズンで到達可能な部分空間上で単射となり、制御に必要な方向を保持できることを証明。 - 従来手法では安定化が不可能な場合があるのに対し、提案手法は制御可能性を保証する点が優れている。

3. 技術・手法の肝は?

- world modelの学習にaction reconstruction objective(逆動力学損失)を追加。 - これによりエンコーダが制御に重要な特徴を保持するcontrol-aware表現を学習。 - 理論的には、正確なaction reconstructionがエンコーダを有限ホライズンHの到達可能部分空間上で単射にすることを証明。 - さらに、Hが大きくなると有限ホライズン可制御性グラミアンの支配固有空間が可制御不安定部分空間に収束することを示す。 - 線形系で理論を構築し、非線形視覚制御タスクに拡張。

4. どうやって有効だと検証した?

- 線形系における理論結果を証明。 - 非線形視覚制御タスク(CartPole, Walker2D, PointMaze)で実験的に検証。 - 提案手法がcontrol-aware表現学習の利点を強調することを示した。

5. 議論はある?

- 次ステップ予測とanti-collapse正則化だけでは制御可能な不安定モードが保持されないことを示し、逆動力学損失の必要性を議論。 - 理論結果は線形系に限定されているが、非線形タスクへの拡張を経験的に示唆。 - 限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- JEPA(Joint-embedding predictive architecture)に関する論文。 - 逆動力学(inverse dynamics)を用いた表現学習の研究。 - 可制御性グラミアン(controllability Gramian)に関する理論。 - 視覚に基づくロボティクス制御のための表現学習手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Leonardo F. Toso, Yann LeCun, James Anderson, Oumayma Bounou

分類: cs.LG, cs.RO, eess.SY, math.OC

原文アブストラクト

Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback. Controlling such systems from high-dimensional visual observations requires representations that preserve these modes. Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning such representations and their dynamics from visual data. However, we demonstrate that next step prediction combined with anti-collapse regularization does not guarantee that controllable unstable modes are preserved: the training loss can be minimized while these modes are collapsed, making stabilization from the learned representation impossible. To address this, we augment world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control. We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace. Thus, the encoder cannot discard any state direction reachable by an action sequence within $H$ steps. Moreover, we show that, as $H$ grows, the dominant eigenspace of the finite-horizon controllability Gramian converges to the controllable unstable subspace. We establish our theoretical results for linear systems and demonstrate empirically that our findings extend to nonlinear visual control tasks (CartPole, Walker2D, and PointMaze), highlighting the benefits of control-aware representation learning.

関連論文

PR本紙発行元 EmplifAI