日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.36985

因果表現学習によるアブダクティブ世界モデリング

Abductive World Modeling via Causal Representation Learning

シェア:XThreadsFacebookLINEはてブBluesky

予測した未来から潜在的な原因をアブダクション推論し、エンティティ・動態・関係の3要素で世界状態を構造化する世界モデルを提案。物理予測や因果推論、行動理解の精度を大幅に向上させた。

詳しい要約

1. どんなもの?

本論文は、世界モデリングにおいて未来状態を予測するだけでなく、その変化を引き起こす潜在的原因を明示的に捉えることを目指す Abductive World Modeling (AWM) を提案する。AWM は予測された未来から潜在的原因を abductive に推論し、構造化された因果表現を学習するフレームワークである。これを Hierarchical Abductive State Pyramid (HASP) で実現し、推論された世界状態を Entity, Dynamic, Relation の3要素に組織化する。これにより、何が存在するか、どのように変化するか、実体間がどう相互作用するかを捉える。

2. 先行研究と比べてどこがすごい?

既存の world models は未来状態を表現するが、その進化の背後にある潜在的原因を明示的に捉えず、なぜ・どのように世界が変化するかの推論が制限されていた。AWM は abductive state inference を latent-space world modeling に初めて導入し、構造化された世界動態表現を学習する点が新しい。実験では V-JEPA と比較して物理予測 AUROC が 10.7%、因果推論精度が 16.8%、行動 Top-1 精度が 68.0% 向上した。

3. 技術・手法の肝は?

AWM の肝は、現在の観測とその予測未来を jointly に推論し、abductive に潜在因子を推定する点にある。これを Hierarchical Abductive State Pyramid (HASP) で実現し、推論された世界状態を Entity, Dynamic, Relation の3つの相補的要素に組織化する。Entity は何が存在するか、Dynamic はどのように変化するか、Relation は実体間の相互作用を捉える。これらを統合して構造化状態表現とし、下流推論に用いる。

4. どうやって有効だと検証した?

物理予測、因果推論、行動理解の実験で有効性を検証した。V-JEPA と比較して、物理予測 AUROC が 10.7%、因果推論精度が 16.8%、行動 Top-1 精度が 68.0% 改善した。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

V-JEPA が比較対象として挙げられている。また、latent-space world modeling や causal representation learning の関連研究が次に読むべき候補として考えられるが、要旨では具体的な参照論文は明示されていない。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ziqi Liu, Songhan Yang, Linfan Zhou, Jiatong Liu, Lijun Peng, Long Wan, Yinqi Bai

分類: cs.LG, cs.AI

原文アブストラクト

The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.

関連論文

PR本紙発行元 EmplifAI