日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.36843

RolloutFaith: 視覚世界モデルにおける持続的内部介入の監査

RolloutFaith: Auditing Persistent Internal Interventions in Visual World Model

シェア:XThreadsFacebookLINEはてブBluesky

世界モデル内部への介入がその後の自律予測にどれだけ持続的な改善をもたらすかを評価する枠組みを提案し、既存の編集手法の限界と遅延学習の有効性を示した。

詳しい要約

1. どんなもの?

- 視覚的 World Model に対する内部介入の持続性を評価する枠組み RolloutFaith を提案。 - 介入時点の予測だけでなく、その後の自律的予測(固定された events, actions, noise, information budgets 下)での意味的改善を測る。 - Crafter, Cartpole, CoinRun 上の3つの world model で10個の fitted editor を評価。 - Reference Activation Patching も用い、選択した interface で得られる補正量を測る。

2. 先行研究と比べてどこがすごい?

- 従来の probe, activation patch, learned editor は現在の計算の解釈・変更に焦点。 - World Model では予測が次の入力になるため、編集停止後も持続する補正が必要という点を指摘。 - 介入時点だけでなく将来の自律予測まで評価する点が新しい。 - 既存の fitted editor は長期効果が限定的・不安定であることを示した。

3. 技術・手法の肝は?

- 固定された events, actions, noise, information budgets の下で、介入時とその後の自律予測の意味的改善を測る。 - Reference Activation Patching: 実観測から計算した paired activation でモデル活性を置換し、interface で利用可能な補正を測る。 - 個々の state component を未介入値に戻し、持続効果の経路を特定。 - Delayed LoReFT: 同じ low rank intervention を4つの凍結された future transition を通して最適化。

4. どうやって有効だと検証した?

- Crafter, Cartpole, CoinRun 上の3つの world model で10個の fitted editor を評価。 - Reference Activation Patching は9つの model-task 組合せすべてで later prediction を改善し、8つで最良の fitted editor を上回った。 - ただし9組合せ中5つで持続的 gain は horizon とともに減少。 - 現行 fitted editor は限定的で一貫しない長期効果のみ。 - 持続効果は DIAMOND では最新生成 frame、DreamerV3 では recurrent memory、STORM では両方を通ると判明。 - Delayed LoReFT は持続的介入効果をある程度改善。

5. 議論はある?

- 持続的効果の経路がモデルにより異なる(DIAMOND: 最新生成 frame、DreamerV3: recurrent memory、STORM: 両方)。 - 訓練時に将来の結果を報酬化すべきという示唆。 - Reference Activation Patching でも horizon とともに gain が減衰する組合せが存在。 - fitted editor の長期効果は限定的・不安定。 - Delayed LoReFT の改善は「ある程度」に留まる。

6. 次に読むべき論文は?

- Reference Activation Patching - DIAMOND - DreamerV3 - STORM - Delayed LoReFT - Crafter, Cartpole, CoinRun 上の world model 関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junchi Yao, Ziyi Wang, Youling Huang, Lijie Hu

分類: cs.LG

原文アブストラクト

Interpretability methods such as probes, activation patches and learned editors are designed to reveal or modify a model's current computation. World models pose a harder requirement: because their predictions become inputs to later predictions, a useful internal correction must survive after editing stops. We therefore propose RolloutFaith, a framework that measures semantic improvement both in the prediction produced at intervention time and over later autonomous predictions under fixed events, actions, noise, and information budgets. We evaluate ten fitted editors on three world models across Crafter, Cartpole, and CoinRun. We also use Reference Activation Patching, which replaces a model activation with the paired activation computed from the real observation, to measure the correction available at the chosen interface. This reference intervention improves later predictions in all nine model and task combinations and outperforms the best fitted editor in eight, yet its sustained gain decreases with horizon in five of nine combinations. Current fitted editors recover only limited and inconsistent long term effects. By restoring individual state components to their untouched values, we find that persistent effects travel through the newest generated frame in DIAMOND, recurrent memory in DreamerV3, and both in STORM. These findings suggest that training should reward future consequences. To test this hypothesis, we propose Delayed LoReFT, which optimizes the same low rank intervention through four frozen future transitions and improves sustained intervention effects to some extent.

関連論文

PR本紙発行元 EmplifAI