日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11794

Memento 3: 反省的ルールブックによるモデルベースの再帰的自己改善

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

シェア:XThreadsFacebookLINEはてブBluesky

凍結したLLMエージェントが自然言語のルールブックを外部記憶として保持し、観察・反省・ルール改訂・コード生成・検証のループを通じて世界モデルを継続的に学習・改善する手法を提案。ARC-AGI-3の全25公開ゲームをクリア。

詳しい要約

1. どんなもの?

Memento 3は、凍結したLLMエージェントが外部メモリを通じて明示的なworld modelを継続的に学習する手法。自然言語のrulebookを永続的なsemantic memoryとして保持し、環境ダイナミクスに関する仮説を記録・改訂する。rulebookを実行可能コードにコンパイルし、予測とplanningに用いる。観察・reflection・rule改訂・コンパイル・検証の継続ループにより、予測誤差を基にrulebookとコードを洗練する。基盤LLMは固定したまま、model-basedなrecursive self-improvement (RSI)の経路を探る。

2. 先行研究と比べてどこがすごい?

Mementoシリーズを基盤に発展。先行研究と比べ、凍結LLMでも外部メモリのrulebookを通じてworld modelを継続学習できる点、rulebookをコードにコンパイルし検証済み更新のみ受け入れる点、population拡張で複数world modelを並列保持し証拠共有する点が特徴。ARC-AGI-3で全25公開ゲームの全レベルをクリアし、mean Relative Human Action Efficiency (RHAE) 100.0、人間の行動数の44%を達成。Atari Pongでは学習済みfeedback controllerがLLM呼び出しなしで3エピソード各21:0勝利。

3. 技術・手法の肝は?

自然言語rulebookをpersistent semantic memoryとして保持し、未知部分は未指定のままにする。rulebookを実行可能コードにコンパイルして予測・planningに使用。観察・reflection・rule改訂・コンパイル・検証の継続ループで予測誤差を利用。更新コードはLLMがrulebookへの忠実性を判断し、cell-exact replayが観察された遷移を再現する場合のみ受理。population拡張では複数world modelを並列保持し、相互作用証拠を共有し予測で探索を誘導。

4. どうやって有効だと検証した?

ARC-AGI-3で単一モデルエージェントが全25公開ゲームの全レベルをクリアし、mean Relative Human Action Efficiency (RHAE) 100.0、人間の行動数の44%を使用。Atari Pongケーススタディでは、学習したfeedback controllerが異なる開始局面の3エピソード各で21:0勝利し、追加のLLM呼び出しなし。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

Mementoシリーズの先行研究、ARC-AGI-3、Atari Pong。関連手法としてrecursive self-improvement (RSI)、world model、external memory、LLM agent。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo, Jun Wang

分類: cs.AI, cs.CL, cs.CV, cs.LG

原文アブストラクト

Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.

関連論文

PR本紙発行元 EmplifAI