日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.02542

世界モデル訓練法:ファインチューニング対RAGの比較

How To Train Your World Model: Fine-tuning vs RAG for LM-based World Modeling

シェア:XThreadsFacebookLINEはてブBluesky

テキストベース環境における言語モデル世界モデル構築で、ファインチューニングとRAGを5環境で比較し、ファインチューニングが優位だがRAGはデータ効率が良いことを示した。

詳しい要約

1. どんなもの?

- テキストベース環境におけるLanguage Model (LM)ベースのWorld Model (WM)の構築手法として、fine-tuningとRetrieval Augmented Generation (RAG)を比較評価した研究。 - 5つの多様な環境(embodied, web navigation, social settings)で体系的に評価。 - fine-tuningが15/20設定で高い報酬を達成する一方、RAGはデータ効率が良いことを示す。 - RAGの検索段階の誤りをcounterfactual interventionで推定し、query reformulation戦略を検討。 - 最終的に、パラメトリックにコアダイナミクスを捉えつつ、アクティブに維持されたmemory storeからの検索に依存するハイブリッドWMを提案。

2. 先行研究と比べてどこがすごい?

- 従来、テキスト環境のWMではfine-tuningが支配的パラダイムであったが、RAGの適用は未探索だった。 - 本研究はfine-tuningとRAGを初めて体系的に比較し、fine-tuningが多くの設定で優位であることを示した。 - また、RAGのデータ効率性やfine-tuningの経験量スケーリングへの依存性を明らかにした。 - さらに、RAGの検索誤りを推定する手法を開発し、query reformulationの階層的アプローチが従来の検索パイプラインを上回ることを示した。 - ハイブリッドWMは複数環境で一貫して他手法を上回る性能を達成。

3. 技術・手法の肝は?

- fine-tuning: LMをテキスト環境の遷移ダイナミクスに適応させる。 - RAG: 経験バッファから関連する遷移を検索し、LMの生成を補助。 - counterfactual intervention: 検索段階の誤り率を推定する手続き。 - query reformulation戦略: 階層的アプローチを含む複数の戦略を検討。 - ハイブリッドWM: パラメトリックにコア環境ダイナミクスを捉え、アクティブに維持されたmemory storeからの検索を学習。

4. どうやって有効だと検証した?

- 5つの多様な環境(embodied, web navigation, social settings)で評価。 - 20の設定でfine-tuningとRAGを比較し、報酬を指標に性能を検証。 - データ効率と経験量スケーリングの影響を分析。 - counterfactual interventionで検索誤りを推定し、query reformulation戦略の有効性を検証。 - ハイブリッドWMを複数環境・モデルで評価し、他手法を上回ることを確認。

5. 議論はある?

- fine-tuningは多くの設定でRAGより優れるが、RAGはデータ効率が良い。 - fine-tuningは経験量のスケーリングから不釣り合いなほど利益を得る。 - RAGの検索段階はしばしば経験バッファから最適でない遷移を表面化させる。 - 階層的query reformulationが従来の検索パイプラインを改善。 - ハイブリッドWMの有効性と頑健性が示唆されるが、更なる検討の余地がある。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: fine-tuning, Retrieval Augmented Generation (RAG), counterfactual intervention, query reformulation, hierarchical approach, hybrid world modelling system。 - 関連手法: Language Model (LM), World Model (WM), experience buffer, memory store。 - 同分野の定番: 強化学習におけるmodel-based planning, テキストベース環境のシミュレーション。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dhananjay Ashok, Shantanu Agarwal, Vivek Datla, Jonathan May, Alfy Samuel

分類: cs.AI

原文アブストラクト

World models (WMs) simulate the transition dynamics of environments, enabling agents to plan over the consequences of their actions. In text-based environments, fine-tuning a Language Model (LM) to serve as a WM has emerged as a dominant paradigm. However, despite the widespread success of non-parametric approaches such as Retrieval Augmented Generation (RAG), retrieval for LM-based world modelling remains underexplored. We conduct a systematic evaluation across five diverse environments spanning embodied, web navigation and social settings, comparing fine-tuning and RAG-based approaches for LM-based world modelling. Our study reveals that fine-tuning often outperforms RAG, with fine-tuned WMs enabling agents to obtain higher rewards on 15/20 settings. While both construction paradigms benefit from additional and more diverse exploration, RAG-based approaches prove more data-efficient, and fine-tuning approaches disproportionately benefit from scaling the amount of experience collected. With a focus on RAG-based WMs, we devise a procedure that uses counterfactual intervention to estimate the error rate of the retrieval stage, and show that retrievers consistently surface suboptimal transitions from the experience buffer. Hoping to address this failing, we study a variety of query reformulation strategies, demonstrating that a hierarchical approach outperforms the traditional retrieval pipeline. Finally, we compose our findings into a hybrid world modelling system that parametrically captures core environment dynamics, while learning to rely on retrieval from an actively maintained memory store. Our hybrid system consistently outperforms other methods across multiple environments and models, showcasing the robustness of the approach and the applicability of our findings.

関連論文

PR本紙発行元 EmplifAI