日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.02204

再構成・練習・実世界展開:身体性エージェントのためのガイド付き自己改善

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

シェア:XThreadsFacebookLINEはてブBluesky

オフラインデータから操作タスクを抽出し、シミュレーションで練習して失敗を診断しながらスキルとプロンプトを自己改善し、実機で高い成功率を達成するフレームワークRPGを提案。

詳しい要約

1. どんなもの?

- ロボット実行システムをモデル重み更新なしで自律改善するフレームワーク RPG (Reconstruct, Practice, Go Real) を提案。 - オフラインデータセットから manipulation capabilities を特定し、simulation 上に関連 practice tasks を構築。 - practice 中に execution feedback、privileged simulator state、dataset videos を用いて失敗を診断。 - 新しい再利用可能な symbolic skills の開発、既存 skills の改良、system prompt の改訂を行う。 - test 時には multimodal LLM が system prompt と skill library を使い perception と robot control を統合。

2. 先行研究と比べてどこがすごい?

- 22 manipulation tasks の held-out initializations で、最初の practice round 後 28.6% から 15 rounds 後 95.0% へ成功率を改善。 - 評価した全 baselines を上回る。ASPIRE は 75.5%、CaP-Agent0 powered by GPT-6 Astra Pro は 60.0%。 - 共通の calibration と hardware-adaptation 手順後、frozen system が 3 tasks 各 10 trials の計 30 physical trials すべてで成功。 - モデル重みを更新せずに自律改善を実現する点が特徴。

3. 技術・手法の肝は?

- offline dataset から manipulation capabilities を特定し、simulation で関連 practice tasks を構築。 - practice 中に execution feedback、privileged simulator state、dataset videos を活用して失敗を診断。 - 診断に基づき、新しい再利用可能な symbolic skills の開発、既存 skills の改良、system prompt の改訂を実施。 - cross-task evaluation により個別候補変更と統合改訂をテストし、保持前に検証。 - test 時は multimodal LLM が system prompt と skill library を用いて perception と robot control を調整。

4. どうやって有効だと検証した?

- 22 manipulation tasks の held-out initializations で評価。 - 成功率が最初の practice round 後 28.6% から 15 rounds 後 95.0% に向上。 - baselines との比較: ASPIRE 75.5%、CaP-Agent0 powered by GPT-6 Astra Pro 60.0% を上回る。 - 共通の calibration と hardware-adaptation 手順後、frozen system が 3 tasks 各 10 trials の計 30 physical trials すべてで成功。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- ASPIRE - CaP-Agent0 - GPT-6 Astra Pro

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi

分類: cs.RO, cs.AI, eess.SY

原文アブストラクト

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/

関連論文

PR本紙発行元 EmplifAI