日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00982

分割して記憶する:長期的VLAポリシーのための再帰的行動関連メモリ

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

シェア:XThreadsFacebookLINEはてブBluesky

現在の観測だけでは行動を決められない履歴依存の操作タスクに向け、行動に関連する情報を履歴から再帰的に選択・記憶する軽量メモリ手法を提案し、長期的操作ベンチマークで最高精度を達成した。

詳しい要約

1. どんなもの?

Vision-language-action (VLA) モデルが苦手とする、現在の観測だけでは行動を決められない履歴依存の manipulation タスク向けの memory 手法。既存手法は「何を覚えるか」を設計者が決めていたが、本論文はそれを最適化問題として捉え、action-relevant な情報を保持する Divide-and-Remember (D&R) を提案。recursive な memory 関数 m_t = M(h_t) を学習し、長い context でも計算軽量に扱える。

2. 先行研究と比べてどこがすごい?

既存の memory 手法は、例えば pixel 変化の大きい frame を保持するなど、何を覚えるかを設計で決めており、タスク間で一貫した利得が得られなかった。D&R は POMDP 定式化から最適な memory が I(a_t; m_t | o_t) を最大化することを示し、これを最適化問題として扱う点が新しい。RoboMME の 16 タスク・4 スイートすべてで一貫した利得を示し、64 token の予算で state-of-the-art の平均成功率を達成。実ロボット実験でも同様の利得を報告。

3. 技術・手法の肝は?

POMDP に基づく imitation learning の解析から、最適な memory は現在の観測に含まれない action-relevant な履歴情報を保持するもの、すなわち I(a_t; m_t | o_t) を最大化するものだと示す。D&R は recursive な memory 関数 m_t = M(h_t) を学習する手法で、(1) 全履歴からの選択を 2K token からの top-K 選択の部分問題へ再帰的に分割し、固定サイズの軽量 selector を end-to-end で学習することで無制限の履歴に対応、(2) すべての recursion block で 1 つの selector を共有し、各 block に共通の選択規則を捉えて効率を保つ。

4. どうやって有効だと検証した?

RoboMME は「いつ・どこで・何を・どのように行動するか」を覚える必要がある 16 の long-horizon manipulation タスクからなる benchmark。D&R は 64 token の予算のみで、4 つのスイートすべてにおいて一貫した利得と state-of-the-art の平均成功率を達成。さらに実ロボット実験でも同じ利得を示した。コード・checkpoint・追加結果は https://dnr-memory.github.io/ で公開。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている具体的な先行 memory 手法名は明記されていない。関連手法として、pixel 変化の大きい frame を保持する設計ベースの memory 手法や、POMDP に基づく imitation learning、VLA モデル、RoboMME benchmark が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh

分類: cs.RO, cs.AI

原文アブストラクト

Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information $I(a_t; m_t \mid o_t)$ between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function $m_t = M(h_t)$ and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-$K$ selection over $2K$ tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at https://dnr-memory.github.io/

関連論文

PR本紙発行元 EmplifAI