日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ワールドモデルarXiv:2610.12016

因果ドリーマー:潜在空間の disentanglement を用いた予測ワールドモデルの学習

CausalDreamer: Learning Predictive World Models with Latent Disentanglement

シェア:XThreadsFacebookLINEはてブBluesky

トークナイザを固定し、その潜在表現を制御可能性と報酬関連性の2軸で4グループに分解して再符号化することで、制御可能・不可能・報酬関連・無関連の情報を分離するワールドモデルを提案。

詳しい要約

1. どんなもの?

- 制御のための世界モデルにおいて、エージェントの行動に応答する要素と報酬に関連する要素を捉えることが重要。 - 生成的世界モデル(例:Dreamer 4)は、ビデオトークナイザーとダイナミクスモデルから成る。 - トークナイザーは再構成目的で訓練され、行動や報酬の監督がないため、潜在表現に制御可能・不能・報酬関連・無関連の情報を分離する仕組みがない。 - 本研究ではCausalDreamerを提案。トークナイザーを凍結し、その潜在表現を2軸(制御性と報酬関連性)に沿って4グループに因子分解する再エンコードを行う。 - 事前訓練済みダイナミクスモデルを因子分解表現予測に微調整する。

2. 先行研究と比べてどこがすごい?

- 先行研究のDreamer 4などの生成的世界モデルは、トークナイザーが再構成目的のみで訓練され、行動や報酬の監督がないため、潜在表現に明示的な分離メカニズムがない。 - CausalDreamerはトークナイザーを凍結し、その潜在表現を制御性と報酬関連性の2軸で4グループに因子分解する点が新しい。 - これにより、制御可能なグループのみが行動を受け取り、報酬関連グループから報酬を予測する。 - 事前訓練済み世界モデルと比較して、クリーンタスクで14%高い正規化スコア(0.199 vs. 0.175)、操作変種で25%高いスコア(0.307 vs. 0.246)を達成。 - 新環境では両モデルともランダムポリシーを有意に上回らないが、因子分解表現が報酬無関連の変化(背景変更など)を報酬関連グループから分離することを示した。

3. 技術・手法の肝は?

- トークナイザーを凍結し、その潜在表現を再エンコードして因子分解表現を得る。 - 因子分解は2軸:制御性(controllability)と報酬関連性(reward relevance)。 - 制御性軸では、4グループのうち制御可能な2グループのみが行動を受け取る。 - 報酬関連性軸では、報酬関連の2グループから報酬を予測するように学習。 - 事前訓練済みダイナミクスモデルを因子分解表現予測に微調整する。

4. どうやって有効だと検証した?

- MMBench2の20タスクでモデル予測計画(model-predictive planning)を評価。 - 内訳:訓練中に見たクリーンタスク10、未見タスク10(うち6は背景・オブジェクト・迷路レイアウトを変更したクリーンタスクの操作変種、4は新環境)。 - リターンを正規化:ランダム行動ポリシーが0、エキスパートが1。 - CausalDreamerは事前訓練済み世界モデルよりクリーンタスクで14%高い正規化スコア(0.199 vs. 0.175)、操作変種で25%高いスコア(0.307 vs. 0.246)。 - 新環境では両モデルともランダムポリシーを有意に上回らない。 - 分析により、因子分解表現が報酬無関連の変化(背景変更など)を報酬関連グループから分離することを示した。

5. 議論はある?

- 新環境では両モデルともランダムポリシーを有意に上回らず、汎化性能に課題が残る。 - 因子分解表現が報酬無関連の変化を分離できることを示したが、その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- Dreamer 4(事前訓練済み世界モデルとして比較) - MMBench2(評価ベンチマーク) - その他、要旨で参照/比較されている研究は明示されていないため、同分野の定番としてDreamerシリーズや世界モデルベースの強化学習手法が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Prince Jha, Nils Lukas, Kun Zhang, Salem Lahlou

分類: cs.LG

原文アブストラクト

World models for control must capture which aspects of the environment respond to the agent's actions and which are relevant to reward. Generative world models such as Dreamer 4 consist of a video tokenizer, which encodes each frame into a latent, and a dynamics model, which is pretrained to predict future latents from past latents and actions. Yet the tokenizer is trained with a reconstruction objective, without action or reward supervision, so its latent provides no explicit mechanism to separate controllable, uncontrollable, reward-relevant, and reward-irrelevant information. We propose \textit{CausalDreamer}, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups. The pretrained dynamics model is then fine-tuned to predict the factored representation. We evaluate \textit{CausalDreamer} and the pretrained world model it starts from with model-predictive planning on 20 MMBench2 tasks: 10 clean tasks seen during training and 10 unseen tasks, of which 6 are manipulated variants of clean tasks with a changed background, object, or maze layout, and 4 are new environments. We normalize returns so that a policy taking uniformly random actions scores 0 and an expert scores 1. \textit{CausalDreamer} achieves a 14\% higher normalized score than the pretrained world model on the clean tasks (0.199 vs.\ 0.175) and a 25\% higher score on the manipulated variants (0.307 vs.\ 0.246), while neither model scores meaningfully above the random policy in the new environments. Additionally, our analysis shows that the factored representation separates reward-irrelevant changes, such as a changed background, from its reward-relevant groups.

関連論文

PR本紙発行元 EmplifAI