日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ロボット自己進化arXiv:2610.12424

RoboRSI:複雑な実世界環境における安定・効率的・再利用可能なロボット自己進化

RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments

シェア:XThreadsFacebookLINEはてブBluesky

タスクをスキル階層に分解し、実行結果を責任スキルに帰属させて検証済みの改善のみを再利用するロボット自己改善システムを提案。実機の家庭内片付け104ラウンドと各種シミュレータで最高性能を達成。

詳しい要約

1. どんなもの?

コードで行動するロボットエージェントが、実行フィードバックからプログラムを修復し、その経験をタスク構造に沿って整理・再利用する自己改善システム RoboRSI を提案する。 - 中核は Top-Down Skill Refinement (TSR) - タスクを compound, atomic, base skills に分解 - 各スキルに責任範囲と明示的な input-output contracts を付与 - 実行結果を責任あるブランチに帰属し、改訂をそのブランチに限定 - Manager, Planner, Engineer, Reviewer が計画・実行・診断・検証済みリリースを調整 - 人間は objectives と corrections で介入 - 安定した skill sequences は再利用可能な compound skills に統合 - mobile manipulator 上で multi-object household cleanup を104ラウンド実施

2. 先行研究と比べてどこがすごい?

実行フィードバックからプログラムを修復する既存のコード行動型ロボットエージェントに対し、経験をタスク構造に基づいて組織化する点が新しい。 - 各修復を責任ある capability に帰属させ、実行証拠で支持し、再利用前に検証する仕組みを導入 - 従来は経験の整理が課題だったが、TSR によりスコープ付き責任と input-output contracts で管理 - シミュレーションで LIBERO, LIBERO-PRO, LIBERO-Plus, RoboTwin において最高成功率を達成し、最強 baseline を 2.7〜11.0 ポイント上回る

3. 技術・手法の肝は?

Top-Down Skill Refinement (TSR) が肝。 - タスクを compound, atomic, base skills に階層分解 - 各スキルに scoped responsibilities と explicit input-output contracts を定義 - 実行結果を責任あるブランチに帰属し、改訂をそのブランチに限定 - Manager, Planner, Engineer, Reviewer の役割分担で計画・実行・診断・検証済みリリースを調整 - 人間が objectives と corrections でプロセスを操舵 - 安定した skill sequences を reusable compound skills に統合

4. どうやって有効だと検証した?

mobile manipulator 上で multi-object household cleanup を104ラウンド実施。 - シミュレーションで LIBERO, LIBERO-PRO, LIBERO-Plus, RoboTwin を評価 - 最高成功率を達成し、最強 baseline を 2.7〜11.0 ポイント上回る - 詳細な実験設定や統計的有意性は要旨からは不明

5. 議論はある?

要旨からは不明。 - 限界、失敗事例、計算コスト、人間介入の負荷、スケーラビリティに関する議論は記述されていない - 実世界での104ラウンドの詳細やシミュレーションとのギャップも不明

6. 次に読むべき論文は?

要旨で参照・比較されている研究は明示されていない。 - 関連手法として LIBERO, LIBERO-PRO, LIBERO-Plus, RoboTwin のベンチマーク論文 - コード行動型ロボットエージェントや program repair を用いたロボット自己改善の研究 - 階層的スキル学習や skill composition の研究が次に読むべき候補

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zimo Wen, Yijin Chen, Yuxuan Cao, Wendi Chen, Yanwen Zou, Wenye Yu, Fuhang Kuang, Han Xue, Jun Lv, Chuan Wen, Cewu Lu

分類: cs.RO, cs.AI

原文アブストラクト

A generalist robot should not only perform diverse tasks but also improve through experience, turning what it learns during execution into capabilities that later tasks can reuse. Robot agents that act through code can already repair programs from execution feedback, yet it remains a central challenge to organize this experience around the task structure that gives it meaning, so that each repair is attributed to the responsible capability, supported by execution evidence, and validated before it is reused. We introduce RoboRSI, a robot self-improvement system built on Top-Down Skill Refinement (TSR). TSR decomposes tasks into compound, atomic, and base skills with scoped responsibilities and explicit input--output contracts, attributes each execution outcome to the responsible branch, and confines revision to that branch. Building upon this structure, a Manager, Planner, Engineer, and Reviewer coordinate planning, execution, diagnosis, and the validated release of new skills, while people steer the process through objectives and corrections; stable skill sequences are further consolidated into reusable compound skills. On a mobile manipulator, RoboRSI develops multi-object household cleanup over 104 rounds. In simulation, it achieves the highest success rate on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin, exceeding the strongest baseline by 2.7 to 11.0 percentage points.

PR本紙発行元 EmplifAI