日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34554

記憶の置き場所:記憶拡張VLAのためのオブジェクト台帳

Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs

シェア:XThreadsFacebookLINEはてブBluesky

短期知覚記憶はポリシー内に、長期物体記憶は外部の読み取り可能な台帳に置くという分割を実現し、単一のファインチューニング済みポリシーでRoboMMEの4スイート平均最高精度を達成した。

著者: Tanguy Dieudonné, Jack B. Jedlicki, Heng Yang

分類: cs.RO

原文アブストラクト

Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned $π_{0.5}$ policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger, built from a SAM3 tracker and a VLM captioner of the demonstration and read by an LLM planner that decides at step boundaries. On RoboMME, Ledger reaches the highest four-suite average among the evaluated methods, 64.3% (vs. 45.9% for the strongest prior method under identical evaluation), leading object reference (60.7% vs. 40.3%) and object permanence (86.7% vs. 56.2%) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.

関連論文

PR本紙発行元 EmplifAI