日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.07028

事前学習済み拡散表現から同定可能な世界モデルを構築

Identifiable World Models from Pretrained Diffusion Representations

シェア:XThreadsFacebookLINEはてブBluesky

凍結した拡散モデルの潜在表現に軽量な整合写像を学習させることで、状態変数や因果構造を同定可能な座標を獲得できることを示した研究。

詳しい要約

1. どんなもの?

拡散モデルベースのworld modelは高次元力学系の軌道を生成・予測できるが、予測精度が高いだけでは潜在座標が真の状態変数や因果相互作用を復元しているとは限らない。本研究は、学習済み拡散モデルを再学習せずに、識別可能な座標を付与できるかを問う。Contrastive Diffusion Alignment (ConDA) を提案し、凍結した拡散潜在表現上に軽量なアラインメント写像のみを学習する。

2. 先行研究と比べてどこがすごい?

従来のworld modelは予測精度を重視し、潜在座標の識別可能性や因果構造の復元を保証しない。また、識別可能な表現学習は通常、生成バックボーンの再学習を要する。本研究は、auxiliary-variable nonlinear ICAの保証をConDAに転用し、凍結拡散モデルでも識別可能・構造解釈可能な座標を得られることを示す点が新しい。

3. 技術・手法の肝は?

凍結したpretrained diffusion modelの潜在表現に対し、軽量なalignment mapのみを学習するContrastive Diffusion Alignment (ConDA) を提案。TCL/GCL仮定下で、整列表現が潜在力学状態を置換と成分ごとの可逆変換まで識別し、潜在動的構造因果モデルを保持し、ラグ付きグラフ復元を遷移ヤコビアンのスパース性に帰着させる。

4. どうやって有効だと検証した?

物理・ロボティクス動画系で、TCL-, GCL-, CEBRA-based ConDAをTDRL, CaRiNG, IDOL, temporal SuaVE, iVAEと比較。TCLとGCLはほぼ完全なブロックワイズ状態復元と競争力のあるラグ付きグラフ復元を達成し、シミュレーション落下物体系では完全復元。シミュレーション二足歩行ロボットでは、学習されたダイナミクスが未見制御摂動への応答の符号と時間構造を復元。

5. 議論はある?

凍結生成拡散モデルに識別可能・構造解釈可能・介入関連ダイナミクス分析に有用な座標を付与できることを示す。ただし、TCL/GCL仮定の成立条件や、実世界ロボットへの適用可能性、他の生成モデルへの一般化については要旨からは不明。

6. 次に読むべき論文は?

TCL, GCL, CEBRA, TDRL, CaRiNG, IDOL, temporal SuaVE, iVAE, auxiliary-variable nonlinear ICA, Contrastive Diffusion Alignment (ConDA) が参照・比較されている。関連分野の定番として、nonlinear ICAやcausal representation learningの基礎文献も次に読むべき。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ruchi Sandilya, Conor Liston, Logan Grosenick

分類: cs.LG

原文アブストラクト

Diffusion-based world models can generate and predict trajectories in high-dimensional dynamical systems, but predictive accuracy does not imply that their latent coordinates recover the underlying state variables or causal interactions. We ask whether a frozen pretrained diffusion model can be equipped with identifiable coordinates without retraining its generative backbone. We show that auxiliary-variable nonlinear ICA guarantees can be transferred to Contrastive Diffusion Alignment (ConDA), which learns only a lightweight alignment map on top of frozen diffusion latents. Under standard TCL/GCL assumptions, the aligned representation identifies latent dynamical states up to permutation and componentwise invertible transformations, preserves the latent dynamic structural causal model, and reduces lagged graph recovery to transition-Jacobian sparsity. We evaluate TCL-, GCL-, and CEBRA-based ConDA against TDRL, CaRiNG, IDOL, temporal SuaVE, and iVAE across physical and robotic video systems. TCL and GCL achieve near-perfect blockwise state recovery and competitive lagged graph recovery, including exact recovery in a simulated falling-body system. In a simulated bipedal robot, learned dynamics recover the sign and temporal structure of responses to held-out control perturbations. These results show that a frozen generative diffusion model can be equipped with coordinates that are identifiable, structurally interpretable, and useful for analyzing intervention-relevant dynamics.

関連論文

PR本紙発行元 EmplifAI