日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
継続学習arXiv:2610.10498

EmbodiedRSI: 仮説駆動型共進化による能動的継続ロボット学習

EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

シェア:XThreadsFacebookLINEはてブBluesky

仮説グラフと情報価値に基づく実験選択でコードとスキルを共進化させ、ロボットの継続学習を効率化する自己進化エージェントを提案。

詳しい要約

1. どんなもの?

- ロボット基盤モデルの性能劣化に対処する自己進化型エージェントハーネス - 仮説グラフでコードとスキルの仮説を維持し、物理実験で検証 - Fast-Slow Dual-System Architectureを採用 - 実験選択とメモリ学習でコード・スキルを共進化 - RoboCasa365やLIBERO-Proで評価、実機にもゼロショット転移

2. 先行研究と比べてどこがすごい?

- 従来の自己進化ハーネスはロボット試行を非効率に使用 - 提案手法はValue-of-Information Experiment Selectionで効率的に仮説を識別 - ベースライン最高40.1%に対し77.0%の成功率 - Composite-Unseenで71.3%を達成 - 実世界タスクでも71.3%のゼロショット成功率

3. 技術・手法の肝は?

- Fast-Slow Dual-System Architecture - Hypothesis Graphで競合するコード・スキル仮説を管理 - Value-of-Information Experiment Selectionで物理実験を選択 - Code-Skill Co-Evolutionで仮説を更新 - Slow SystemがHierarchical Memoryを構築 - Reward-Grounded Memory Learningで有効なメモリを選択

4. どうやって有効だと検証した?

- RoboCasa365で全体成功率77.0%、Composite-Unseenで71.3% - ベースライン最高40.1%と比較 - LIBERO-Proで全体成功率86.8% - 実世界ロボットにゼロショット転移し、複数タスクで71.3%の成功率

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- RoboCasa365 - LIBERO-Pro - ロボット基盤モデル - 自己進化ハーネス - テレオペレーション

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Python Song, Zhixuan Liang, Kelsey Fu, Mengdi Wang, Junfeng Yang, Shilong Liu

分類: cs.AI

原文アブストラクト

Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-of-Information Experiment Selection chooses physical experiments that can distinguish these hypotheses. Their outcomes guide Code-Skill Co-Evolution. The Slow System builds Hierarchical Memory, and Reward-Grounded Memory Learning selects effective memory according to their value for later Fast-System improvement. On RoboCasa365, EmbodiedRSI reaches 77.0% overall success and 71.3% on Composite-Unseen, compared with 40.1% for the best baseline. EmbodiedRSI also reaches 86.8% overall success on LIBERO-Pro. Beyond benchmark performance, EmbodiedRSI transfers zero-shot to real-world robot, achieving 71.3% overall success across multiple challenging tasks.

関連論文

PR本紙発行元 EmplifAI