MIKASA-Robo-VLA:長期的マニピュレーションにおけるVLAモデルの記憶能力を評価するベンチマーク
MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
言語条件付きマニピュレーション90タスクからなるベンチマークを構築し、タスク遂行に必要な手がかりを隠すことでVLAモデルの記憶能力を評価する。
著者: Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev
分類: cs.LG, cs.AI, cs.RO
原文アブストラクト
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference $π_{0.5}$ baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 $\pm$ 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/