再構成・練習・実世界展開:身体性エージェントのためのガイド付き自己改善
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
オフラインデータから操作タスクを抽出し、シミュレーションで練習して失敗を診断しながらスキルとプロンプトを自己改善し、実機で高い成功率を達成するフレームワークRPGを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi
分類: cs.RO, cs.AI, eess.SY
原文アブストラクト
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
関連論文
- スクリューアテンション:Transformer内部に剛体代数を組み込むマニピュレーション
- SkeleWAM: 骨格ワールドアクションモデリングによる効率的なロボットマニピュレーションマニピュレーション
- FlashDexRetarget: 多動作リターゲティングによる器用操作データ生成の高速化マニピュレーション
- 解像度に一貫したヤコビアン場を学習する生体模倣剛柔指マニピュレーション
- Recova: 自律ロボットマニピュレーションのためのエージェント誘導型失敗回復マニピュレーション
- 経験と実演による6自由度把持合成の継続学習マニピュレーション