日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
継続学習arXiv:2610.03079

RIFAR: 信頼性と忘却を考慮した継続的ロボット学習のためのリプレイ手法

RIFAR: Reliability and Forgetting-Aware Replay for Continual Robot Learning

シェア:XThreadsFacebookLINEはてブBluesky

生成的な世界行動モデルによるリプレイで、行動と観測の整合性を評価して高品質な軌道を選別し、適応後の予測ドリフトに基づいて履歴を再選択することで、継続学習での忘却を抑える手法を提案。

詳しい要約

1. どんなもの?

- 継続的なロボット学習において、新しいタスクを学びながら過去の知識を忘れないための手法「RIFAR」を提案。 - Experience replayの代替として、World-action models (WAM) を用いた生成的リプレイを採用。 - 信頼性スクリーニングとドリフト対応リプレイ選択を組み合わせる。 - タスクと環境が進化する中で、転移可能なスキルを維持することを目指す。

2. 先行研究と比べてどこがすごい?

- 従来のExperience replayは完全なデモンストレーションの保存がコスト高。 - WAMベースの生成的リプレイは視覚的に一貫したロールアウトでも、予測遷移を実現できない行動を含む可能性がある。 - 新タスク適応が以前の学習行動を破壊することがある。 - RIFARは信頼性スクリーニングとドリフト対応選択でこれらの問題に対処し、LIBERO-Goalで90.97 AUCを達成、50デモリプレイの約4.9%のステップ数で過去の状態を保持。

3. 技術・手法の肝は?

- コンパクトなデモンストレーションのプレフィックスから軌道を再構築。 - 凍結したinverse-dynamics modelを用いて行動と視覚の一貫性を評価。 - 訓練は現在のデモンストレーションと最高品質のスクリーニング済み軌道を組み合わせる。 - 適応前後で同一の履歴入力に対する行動予測を比較し、正規化ドリフトが大きい軌道を同じスクリーニングプールから再選択して継続訓練。

4. どうやって有効だと検証した?

- 3つのLIBEROスイートと実世界実験で評価。 - WAMベースの生成的リプレイにおける従来のstate of the artを上回る。 - LIBERO-Goalで90.97 AUCを達成し、タスクあたり320の履歴タイムステップのみを保持(50デモリプレイの約4.9%)。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- LIBEROスイート(LIBERO-Goalなど) - World-action models (WAM) - Experience replay - Inverse-dynamics model

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xiuying Chen

分類: cs.AI

原文アブストラクト

Genuine embodied agency requires robots to turn continuous real-world experience into lasting, transferable skills. This demands continual learning that integrates new capabilities without eroding prior knowledge as tasks and environments evolve. Experience replay mitigates forgetting, but storing complete demonstrations becomes costly as tasks accumulate. World-action models offer a generative alternative, reconstructing past experience through joint predictions of actions and future observations. However, visually coherent rollouts may contain actions that cannot realize the predicted transitions, while new-task adaptation can disrupt previously learned behavior. RIFAR therefore combines reliability screening with drift-aware replay selection. It reconstructs trajectories from compact demonstration prefixes and uses a frozen inverse-dynamics model to assess action-visual consistency. Training first combines current demonstrations with the highest-quality screened trajectories. RIFAR then compares action predictions before and after this adaptation on identical historical inputs, reselecting trajectories with larger normalized drift from the same screened pool for continued training. Across three LIBERO suites and real-world experiments, RIFAR surpasses the previous state of the art in WAM-based generative replay. On LIBERO-Goal, it achieves 90.97 AUC while retaining only 320 historical time steps per task, approximately 4.9% of the steps retained using 50-demonstration replay.

関連論文

PR本紙発行元 EmplifAI