日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
身体性記憶/ベンチマークarXiv:2609.28236

EmbodiedMemory-Bench:長期的な身体性タスクにおける身体性記憶のベンチマーク

EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

シェア:XThreadsFacebookLINEはてブBluesky

長期的な身体性インタラクションにおける記憶能力を評価するEMem-Benchを提案し、空間・イベント・シーン記憶を統合する外部記憶システムEMemと8BポリシーEMem-8Bを開発した。

詳しい要約

1. どんなもの?

- 長期的な embodied interaction における memory 能力を評価する benchmark と memory system の提案。 - EmbodiedMemory-Bench (EMem-Bench): 4 つの task family にわたる 2,554 の interactive episode を含む。 - エージェントは interaction history から memory を構築・更新し、後の task を環境内で行動して完了する必要がある。 - Embodied-Memorizer (EMem): embodied experience を spatial, event, scene memory に組織化する external memory system。 - EMem-8B: 8B policy で memory を管理・利用するよう訓練。

2. 先行研究と比べてどこがすごい?

- 既存 benchmark は long-horizon embodied interaction 中の memory 能力を直接評価していない。 - 著者らの分析で、現在のエージェントの限界を 4 つの欠陥に特定: 弱い fine-grained visual memory、信頼できない dynamic world-state tracking、interaction outcome が明かす world state の記録失敗、prior experience からの限定的な generalization。 - EMem-Bench はこれらの memory 能力を直接評価する点で先行 benchmark と異なる。 - EMem は評価された memory system の中で matched backbone 下で最高の overall performance を達成。 - open-source と proprietary の両モデルを改善し、EMem-8B は backbone をさらに上回る。

3. 技術・手法の肝は?

- EMem-Bench: 4 つの task family にわたる 2,554 の interactive episode で構成。 - エージェントは interaction history から memory を構築・更新し、後の task で環境内行動に利用。 - EMem: embodied experience を spatial, event, scene memory に組織化する external memory system。 - EMem-8B: 8B policy で memory を管理・利用するよう訓練。 - 多様な open-source および proprietary MLLM と代表的な multimodal memory system を評価。

4. どうやって有効だと検証した?

- 多様な open-source および proprietary MLLM と代表的な multimodal memory system を評価。 - 結果、現在のモデルは 4 つの challenge で弱く不均一。 - matched backbone 下で EMem は評価された memory system 中最高の overall performance。 - EMem は open-source と proprietary モデルの両方を改善。 - EMem-8B は backbone をさらに改善。

5. 議論はある?

- 現在のエージェントは long-horizon embodied interaction で memory を信頼性高く維持できない。 - 4 つの欠陥: 弱い fine-grained visual memory、信頼できない dynamic world-state tracking、interaction outcome が明かす world state の記録失敗、prior experience からの限定的な generalization。 - 既存 benchmark はこれらの memory 能力を直接評価しない。 - 現在のモデルは 4 つの challenge で弱く不均一。 - その他の議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 既存 benchmark、代表的な multimodal memory system。 - 関連手法: MLLM、multimodal memory system、embodied interaction の long-horizon task。 - 具体的な論文名は要旨からは不明。 - 同分野の定番: embodied AI、long-horizon planning、memory-augmented agents に関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lizhou Liang, Xinyu Zhong, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Qinfeng Li, Peng Li, Jintao Chen, Xuhong Zhang, Wenqi Zhang

分類: cs.CV

原文アブストラクト

Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, unreliable dynamic world-state tracking, failing to record world state revealed by interaction outcomes, and limited generalization from prior experience. However, existing benchmarks do not directly assess these memory capabilities during long-horizon embodied interaction. To address this gap, we introduce EmbodiedMemory-Bench (EMem-Bench), comprising 2,554 interactive episodes across four task families. EMem-Bench requires agents to build and update memory from interaction history, then use it to complete a later task by acting in the environment. We further present Embodied-Memorizer (EMem), an external memory system that organizes embodied experience into spatial, event, and scene memories. We also train EMem-8B, an 8B policy that manages and uses these memories. We evaluate a diverse range of open-source and proprietary MLLMs and representative multimodal memory systems. Results show that current models remain weak and uneven across the four challenges. Under matched backbones, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models, while EMem-8B further improves over its backbone. Project page: https://zju-omniai.github.io/EmbodiedMemoryBench/

PR本紙発行元 EmplifAI