日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.04701

ロボット方策のための残差メモリとしてのテスト時訓練

Test-Time Training as Residual Memory for Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

テスト時訓練(TTT)を「残差メモリ」として再解釈し、現在の観測と限られたメモリバンクから復元できない履歴情報だけを高速重みに保持することで、長期的なロボットマニピュレーションの記憶問題を効率化する手法TTT-RMを提案。

詳しい要約

1. どんなもの?

- 長期的なロボット操作における記憶問題を扱う研究 - Test-Time Training (TTT) を Residual Memory として再解釈する TTT-RM を提案 - 限られた memory bank と fast weights を組み合わせ、現在の観測から復元できない履歴情報を保持 - シミュレーションと実世界タスクで検証

2. 先行研究と比べてどこがすごい?

- 従来は過去観測を bounded memory bank に保存するか、TTT で履歴を固定サイズのパラメトリック状態に圧縮 - 従来手法は現在の文脈から復元可能な情報と、そうでない情報を明示的に区別しない - TTT-RM は TTT を汎用的な圧縮ではなく、bounded memory bank を補完する residual memory として利用 - 現在の文脈から復元できないタスク関連履歴のみを保持

3. 技術・手法の肝は?

- history decoder を学習し、現在の観測と取得した memory から過去の表現を再構成 - 再構成残差 (reconstruction residual) を計算し、これが現在の文脈で説明できない情報を捉える - この残差を TTT の学習ターゲットとする - TTT の slow weights を残差ターゲットに対して最適化し、オンラインの fast-weight 更新が補完的な履歴情報を符号化 - fast-weight 状態をクエリして residual memory 表現を生成し、行動生成の条件付けに使用

4. どうやって有効だと検証した?

- 記憶集約型のシミュレーションベンチマークと実世界タスクで広範な実験 - 複数の memory 設計に対して一貫した改善を確認 - 多様なベースラインを上回る性能 - 3分間・8段階の stowing タスクで持続的な実行をサポート

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- Test-Time Training (TTT) を用いた記憶手法 - bounded memory bank を利用する既存研究 - 長期的ロボット操作のための memory-augmented policies - 具体的な論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoxuan Wang, Gengyu Zhang, Ramana Rao Kompella, Gaowen Liu, Yan Yan

分類: cs.RO

原文アブストラクト

Memory is essential for long-horizon robotic manipulation, where successful actions may depend on past events that are no longer recoverable from the current observation. As episodes grow longer, however, retaining the full history becomes increasingly costly, creating a fundamental scalability challenge for memory-augmented policies. Existing approaches address this challenge by either storing selected past observations in a bounded memory bank or compressing interaction history into a fixed-size parametric state through Test-Time Training (TTT). Yet these formulations do not explicitly distinguish between historical information that can already be recovered from the policy's current context and information that must persist beyond it. We introduce TTT-RM, which repurposes TTT as Residual Memory, using fast weights not to generically compress history but to complement a bounded memory bank by preserving task-relevant historical information that cannot be recovered from the policy's current context. Concretely, TTT-RM learns a history decoder that reconstructs historical representations from the current observation and retrieved memory. The resulting reconstruction residual captures what this context fails to explain and serves as the learning target for TTT. The TTT slow weights are optimized against this residual target so that online fast-weight updates learn to encode complementary historical information over time. The fast-weight state is then queried to produce a residual memory representation that conditions action generation. Extensive experiments on memory-intensive simulation benchmarks and real-world tasks show that TTT-RM consistently improves across multiple memory designs, outperforms diverse baselines, and supports sustained execution on a three-minute, eight-stage stowing task.

関連論文

PR本紙発行元 EmplifAI