日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.37378

Do-JEPA: 潜在世界モデルにおけるマスキングから介入へ

Do-JEPA: From Masking to Intervention in Latent World Models

シェア:XThreadsFacebookLINEはてブBluesky

行動による効果と単なる共起を分離するため、参照行動との潜在的未来の差分を予測する介入ベースの世界モデル学習法を提案し、効果の伝播や不変性を捉える損失で精度を向上させた。

詳しい要約

1. どんなもの?

- 潜在世界モデル(latent world model)の学習目的を、単なる次状態予測から「行動が引き起こした効果」の予測へ変える手法 Do-JEPA を提案。 - 同一の保存済みシミュレータ状態から、行動 a と参照行動 a_∅ の2つの動力学を走らせ、その潜在未来の差 Δz = z^a − z^{a_∅} を予測するよう学習。 - 目的関数は effect loss、support loss(行動がどこに入るか)、propagation loss(効果がどこへ伝播するか)、invariance losses(変化してはいけないもの)から構成。 - 対象は object-aligned 変数を持つ合成系からピクセル入力まで。

2. 先行研究と比べてどこがすごい?

- C-JEPA など object-masking 系は「予測器が見えるもの」に介入するのに対し、Do-JEPA は「物理的に起きること」に介入する点が異なる。 - 合成系で support supervision は直接介入された object をテストケースの 99.95% で見つけるが、sparse action mask は全ケースで nuisance slot に行動を送ってしまう。 - response-onset supervision は ring 型の propagation graph を回復(edge AUROC 0.975 対 0.624)。 - ピクセル入力では effect loss が同一データで訓練した control を上回る。

3. 技術・手法の肝は?

- 1つの保存済みシミュレータ状態から、行動 a と参照行動 a_∅ の下での動力学を実行し、2つの潜在未来の差 Δz を教師信号として予測させる。 - 目的関数は effect loss、support loss、propagation loss、invariance losses の4要素。 - support loss は行動がどこに入るかを、propagation loss はその効果がどこへ伝播するかを、invariance losses は変化してはいけないものを規定。 - ピクセルから end-to-end の LeWM モデルにも適用可能。

4. どうやって有効だと検証した?

- 合成系(object-aligned 変数)で support supervision と response-onset supervision の性能を評価。 - ピクセル入力で effect loss を同一データの control と比較:end-to-end LeWM で latent effect error を 28.4% 低減、自然な行動系列で訓練・テストした場合に physical effect error を 13.5% 低減。 - 独立生成した3つの CausalWorld ベンチマークで、physics shift 下の responsive effect error を約20%低減、予測効果の latent context sensitivity を66%低減。 - スクラッチ訓練では factual accuracy を損なうが、既存モデルの fine-tuning ではこのコストが除去される。

5. 議論はある?

- スクラッチから訓練すると factual accuracy にコストが生じるが、既存モデルを fine-tuning するとこのコストは除去される。 - 世界に介入することが、モデルが見るものに介入するより、行動が引き起こすことを予測するのに役立つことを示唆。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- C-JEPA(object-masking 系の先行研究) - LeWM(end-to-end の latent world model) - CausalWorld(ベンチマーク) - 関連する latent world model や JEPA 系の研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hossein Resani, Javen Qinfeng Shi

分類: cs.CV, cs.AI, cs.LG

原文アブストラクト

Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $Δz=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.

関連論文

PR本紙発行元 EmplifAI