日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ワールドモデルarXiv:2609.37441

異方性表現でJEPAワールドモデルの計画を改善

Anisotropic Representations Improve Planning in JEPA World Models

シェア:XThreadsFacebookLINEはてブBluesky

JEPA型ワールドモデルの潜在空間正則化を等方ガウスから学習可能な対角共分散に置き換え、計画コストとタスク整合性を高めるAnisoWMを提案。4つの視覚制御環境で計画成功率を改善。

詳しい要約

1. どんなもの?

- JEPA 系の latent world model における表現幾何と planning の整合性を扱う研究。 - 行動条件付きダイナミクスを表現空間で学習し、goal 表現への Euclidean 距離で候補行動を評価する枠組みが対象。 - 等方 Gaussian 正則化が collapse を防いでも、task cost と異なる順位付けを生む問題を指摘。 - 学習可能な対角共分散 Λ を導入する AnisoWM と ΛReg を提案。 - 予測目的・predictor 構造・Euclidean planner は変えず、学習時のみ target を変更。

2. 先行研究と比べてどこがすごい?

- 従来は collapse 防止の正則化と予測精度が planning 性能を保証すると暗黙に想定。 - 本研究は等方 Gaussian 正則化が task-aligned でない幾何を誘導し得ることを示す。 - 固定等方 Gaussian target を、固定 trace と異方性制約下の学習可能対角共分散に置換。 - 予測器や planner を変更せず、表現幾何のみを調整する点が差分。 - 4 つの visual control 環境で LeWorldModel より planning 成功率が向上。

3. 技術・手法の肝は?

- AnisoWM は target 分布を学習可能な対角共分散 Λ に置き換える ΛReg を採用。 - 固定 trace 制約と異方性制約を課し、collapse を防ぎつつ方向ごとの分散を調整。 - 予測目的・predictor architecture・Euclidean planner は不変で、target は学習時のみ使用。 - 予測駆動的な target variance の配分を解析し、訓練分布への依存を特徴づける。 - 誘導される metric が planning regret を低減する条件を理論的に示す。

4. どうやって有効だと検証した?

- 4 つの visual control 環境で評価。 - AnisoWM は全 4 環境で LeWorldModel より planning 成功率が向上。 - latent planning cost が task outcome とより一致することを確認。 - 予測駆動的な target variance 配分と訓練分布依存性を解析。 - planning regret 低減条件を理論的に特徴づけ。

5. 議論はある?

- 正確な予測と非 collapse 表現だけでは task-aligned な latent planning cost を保証しない。 - 等方 Gaussian 正則化は feasible outcome の順位を task cost と異ならせ得る。 - 提案 metric が planning regret を低減する条件と訓練分布依存性を議論。 - 予測器や planner を変えず target のみ変更する設計の含意を議論。 - 詳細な限界や失敗条件は要旨からは不明。

6. 次に読むべき論文は?

- LeWorldModel(比較対象の latent world model)。 - JEPA 系 world model の代表的研究(I-JEPA, V-JEPA など)。 - latent world model における collapse 防止正則化(VICReg, Barlow Twins など)。 - model-based planning における latent cost 設計に関する研究。 - visual control ベンチマーク(DMControl など)を用いた planning 評価研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mingu Kang, Yoori Oh, Sookyung Kim, Joonseok Lee

分類: cs.RO

原文アブストラクト

Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with $Λ$Reg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/

関連論文

PR本紙発行元 EmplifAI