日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
記憶/計画arXiv:2607.14252v2

MEMORA: 一人称視点ビデオからの身体化された行動記憶による推論と計画

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

シェア:XThreadsFacebookLINEはてブBluesky

エージェントの経験を永続的な記憶として形成・利用する枠組みMEMORAを提案し、45時間の一人称視点ビデオベンチマークで計画性能を最大16.6%向上させた。

著者: Zihao Yu, Xiu Yuan, Chongjie Zhang

分類: cs.RO, cs.AI, cs.CL

原文アブストラクト

Embodied agents accumulate experience over time. We study how accumulated experience can be formed into persistent memory for future reasoning and action. We formulate Embodied Action Memory (EAM) as the capability to form and use memory over embodied experience, together with the persistent memory state produced by that process. We introduce MEMORA, a framework that instantiates EAM through a formation-consolidation-retrieval lifecycle and a multi-store world-memory architecture. MEMORA organizes experience into participant-specific Environment, Entity, Activity, and Inferred Knowledge stores: online editing revises memory as new evidence arrives, while offline consolidation abstracts repeated experience into reusable routines, habits, and preferences. We evaluate MEMORA with MEMORA-Bench, a 45-hour egocentric-video suite that measures both retrospective memory faithfulness and prospective memory-grounded planning. Across four open-weight answer models, MEMORA achieves the strongest aggregate planning performance among the evaluated memory interfaces, with its largest gains on out-of-distribution planning. On these tasks, MEMORA improves Robot-Grounded Plan score by up to 16.6 percent, suggesting that memory formed and consolidated across experience can support planning for new goals beyond directly observed episodes. A physical-robot demonstration further shows that memory formed solely from human egocentric video can ground high-level robot plans in participant-specific objects and preferences. Project website: https://github.com/yuzihaowashu/MEMORA