日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.34058

世界モデルは大域的理解を学習するのか?

Do World Models Learn Global Understanding?

シェア:XThreadsFacebookLINEはてブBluesky

モノイド世界での学習課題を通じ、局所的な遷移から大域的制約を学習・伝播できるかを検証。合成訓練により逆元・可換・合成制約で96%の精度を達成し、世界モデルやLLMの汎化も改善することを示した。

詳しい要約

1. どんなもの?

- 世界モデルやLLMが局所情報から大域的理解を学習できるかを調査。 - monoid worlds(状態と行動遷移の集合)上で、訓練遷移と未見制約から未見遷移を予測するタスクを構築。 - 逆・可換・合成・周期性の制約を扱い、一般化性能を測定。 - 次状態予測訓練は非自明な制約伝播に失敗するが、合成訓練は高精度を達成。 - 証明深度dの増加に伴い一般化が急減することを発見。

2. 先行研究と比べてどこがすごい?

- 従来の世界モデルやLLMは局所情報を大域的理解に持ち上げるのが苦手とされる。 - 本研究は「理解」を制約学習とその帰結の伝播として形式化し、一般化を測定する枠組みを提案。 - 次状態予測訓練では非自明な制約伝播が失敗することを示し、合成訓練が有効であることを発見。 - 幾何的一般化や関係的一般化の改善も確認し、先行研究との比較で優位性を示す。

3. 技術・手法の肝は?

- monoid worldsを構築し、逆・可換・合成・周期性の制約を設定。 - 訓練遷移と未見制約から未見遷移を予測するタスクを設計。 - 次状態予測訓練と合成訓練(同一経路だが中間状態を隠す)を比較。 - 証明深度dを定義し、推論ラウンドの最小数を測定。 - 合成経路長Tを増加させ、一般化への影響を評価。

4. どうやって有効だと検証した?

- attention, recurrent, state-spaceアーキテクチャで実験。 - 合成訓練が逆・可換・合成制約で96%の精度を達成。 - 身体性環境で訓練した世界モデルの幾何的一般化と、Wikidataで微調整したLLMの関係的一般化の改善を確認。 - 証明深度dの増加に伴う一般化の急減を定量的に示す。

5. 議論はある?

- モデルが未見事実を推論する際、他の事実を先に推論する必要がある場合、制約伝播がどの程度進むかは証明深度に依存。 - 合成訓練が情報伝播と統合を促進することを実証。 - 大域的理解を形式化する方法を提供し、言語モデルと世界モデルの研究に貢献。 - 限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、world models, LLMs, attention, recurrent, state-space architectures, Wikidata-finetuned LLMsが挙げられる。 - 同分野の定番として、model-based reinforcement learningやtransformer-based world modelsが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Alexander Detkov, Matt Thomson

分類: cs.LG, cs.AI

原文アブストラクト

AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we frame "understanding" as learning constraints and propagating their consequences. We construct learning tasks on monoid worlds, sets of states connected by action transitions, where observed training transitions and an unseen constraint jointly determine held-out transitions. Measuring generalization tests whether models can learn global constraints from local transitions and propagate their consequences. We consider inverse, commutativity, composition, and periodicity constraints relevant to spatial and semantic structure. Across attention, recurrent, and state-space architectures, next-state training fits the data but fails to propagate non-trivial constraints. Compositional training, which uses identical paths but hides intermediate states from the input, achieves 96% accuracy on inverse, commutativity, and composition constraints across architectures, yields corresponding improvements in geometric generalization of world models trained on embodied environments and relational generalization in Wikidata-finetuned LLMs. How far do models propagate constraints when inferring an unseen fact may depend on first inferring others? We define proof depth d of a held-out transition, measuring the minimum number of inference rounds to infer the transition, and find that model generalization decreases sharply with proof depth. Increasing compositional path length T improves generalization. These results provide a formal way to investigate global understanding in language and world models and demonstrate that compositional training promotes information propagation and integration.

関連論文

PR本紙発行元 EmplifAI