日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.35200

ReCAT: 記憶・計数・時間を考慮した構造化リカレントメモリによるロボットマニピュレーション

ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

言語条件付きポリシーにMamba-2と因果注意層による構造化リカレントメモリを組み込み、過去の視覚手がかりの想起やイベント計数、経過時間推定を必要とするロボット操作タスクを高精度で実現した。

詳しい要約

1. どんなもの?

- 言語条件付きロボットマニピュレーションのための構造化リカレントメモリを持つポリシーReCATを提案。 - メモリ依存タスク(視覚的手がかりの想起、タスク進捗追跡、イベント計数、経過時間推定)を対象。 - 命令条件付きエンコーダ、Mamba-2層と因果注意層からなるリカレントメモリ、フローマッチングTransformerデコーダで構成。 - LIBEROで95.3%、RMBenchで62.4%の平均成功率を達成。 - 実機タスク(空間想起、イベント計数、間隔タイミング)で最高66.7%の成功率。

2. 先行研究と比べてどこがすごい?

- 従来の短履歴ベースライン(最強で8.3%)に対し、実機タスクで66.7%と大幅に優れる。 - LIBEROとRMBenchで9タスク中6タスクで最高または同等の結果。 - メモリ更新規則(加算更新、デルタ規則更新)がロボットメモリとして異なる振る舞いを示すことを明らかにした。 - 観測エンコーダと毎ブロックのメモリ条件付けが性能に必要であることを制御比較で示した。

3. 技術・手法の肝は?

- 命令条件付きエンコーダが現在の観測から特徴を形成。 - リカレントメモリがMamba-2層と1つの因果注意層を通じて観測ストリームを統合。 - フローマッチングTransformerデコーダが各ブロックで別々のクロスアテンションを介して現在と履歴の表現を読み取る。 - メモリ更新規則として加算更新とデルタ規則更新を比較。

4. どうやって有効だと検証した?

- LIBEROとRMBenchのベンチマークで評価。 - 実機で空間想起、イベント計数、間隔タイミングの3タスクを実施。 - 制御比較により、観測エンコーダと毎ブロックのメモリ条件付けの必要性を検証。 - メモリ更新規則の違いによる成功率の変化を観察。

5. 議論はある?

- メモリ更新規則がロボットメモリとして異なる効果を示す:加算更新は計数とタイミングで最高、デルタ規則更新は空間想起で最高。 - 観測エンコーダと毎ブロックのメモリ条件付けが性能に必要。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- Mamba-2 - フローマッチングTransformer - LIBERO - RMBench - デルタ規則更新 - 因果注意層

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj, Zhuoyue Li, Moritz Reuss, Rudolf Lioutikov

分類: cs.RO, cs.AI, cs.CV, cs.LG

原文アブストラクト

Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3\% average success on LIBERO and 62.4\% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7\% average success, against 8.3\% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT

関連論文

PR本紙発行元 EmplifAI