日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
エージェント記憶arXiv:2610.10778

世界モデルがエージェントの記憶を変える:MemoWM

MemoWM: How World Models Change What Agents Need to Remember

シェア:XThreadsFacebookLINEはてブBluesky

世界モデルの予測を活用して長期エージェントの記憶を圧縮・再構成する枠組みを提案し、精度を向上させつつ保存量を大幅削減した。

詳しい要約

1. どんなもの?

- 長期エージェントの記憶をWorld Modelで圧縮する枠組みMemoWMを提案 - World Modelが捉える再利用可能な規則性を利用し、経験ごとの保存情報を削減 - 共有予測で保持情報を圧縮し、省略内容を再構成する - 記憶割当をWorld Model条件付き問題として定式化 - 5つの長期エージェント記憶ベンチマークで評価

2. 先行研究と比べてどこがすごい?

- 最も記憶効率の良いbaseline MIRIXに対し、経験固有storageを平均53.9%削減 - 平均回答精度42.42%で最強baselineを2.62ポイント上回る - 従来の記憶圧縮と異なり、World Modelの予測priorを記憶割当に活用 - 強いWorld Modelほど同一タスク品質で経験あたりstorageを削減できることを示す - モデルパラメータも含めた総storageのtrade-offを分析

3. 技術・手法の肝は?

- World Modelの共有予測を用い、保持情報の圧縮と省略内容の再構成を行う - task-aware allocation ruleを導入 - 再構成誤差の期待影響とstorage costのバランスを取る - 予測priorを超える下流価値を持つ情報を保持 - 共有モデル容量と反復storageコストのtrade-offを考慮

4. どうやって有効だと検証した?

- 5つの長期エージェント記憶ベンチマークで評価 - 平均回答精度42.42%、最強baselineを2.62ポイント上回る - MIRIX比で経験固有storageを平均53.9%削減 - 強いWorld Modelが経験あたりstorageを削減することを分析で確認 - 保持interaction数が増えると総storage最小化容量が増えることを示す

5. 議論はある?

- 強いWorld Modelは同一タスク品質で経験あたりstorageを削減 - モデルパラメータを含めると共有モデル容量と反復storageコストにtrade-off - 総storageを最小化する容量は保持interaction数が増えると増加 - 再構成誤差の下流影響とstorage costのバランスが重要 - 具体的な限界や失敗事例は要旨からは不明

6. 次に読むべき論文は?

- MIRIX(最もstorage効率の良いbaseline) - World Model関連の長期エージェント記憶研究 - 長期エージェント記憶ベンチマーク群 - 記憶圧縮・再構成手法 - コード: https://github.com/Feld-maxiu/MemoWM

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bingfan Zeng, Zhisheng Chen, Chenbo Sang, Zhengwei Xie, Jinpeng Wang, Xiangchen Guan, Rui Qian, Zheng Lu, Jingwei Song

分類: cs.LG, cs.AI

原文アブストラクト

Long-term agents face growing storage demands as they accumulate experience. World models capture reusable regularities that can reduce the information stored for each experience. We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content. Its task-aware allocation rule balances the expected impact of reconstruction errors against storage cost, retaining information with downstream value beyond the predictive prior. Across five long-term agent-memory benchmarks, MemoWM achieves 42.42\% average answer accuracy, exceeding the strongest baseline by 2.62 percentage points, while reducing average experience-specific storage by 53.9\% relative to MIRIX, the most storage-efficient baseline. Further analysis shows that stronger world models reduce per-experience storage at comparable task quality. Accounting for model parameters reveals a trade-off between shared model capacity and recurring storage costs, with the capacity that minimizes total storage increasing as more interactions are retained. Our code is available at https://github.com/Feld-maxiu/MemoWM.

関連論文

PR本紙発行元 EmplifAI