混合品質の実環境経験からロボットマニピュレーションを学習する予測的行動チャンク学習
Learning from Mixed-Quality Deployment Experience for Robot Manipulation
実環境で蓄積された成功・失敗・部分進捗が混在するロボットの経験を、チャンク単位の予測的クリティックと拡散方策で活用し、追加の人間修正なしに方策を改善する手法PACLを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yangang Ren, Yujie Yan, Zirui Li, Jiaming Guo, Di Zeng, Ji Tao, Lan Yu, Xuesong Tian, Chen Lv
分類: cs.LG
原文アブストラクト
Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.
関連論文
- Streaming-WAM: 非同期ロボットマニピュレーションのための行動条件付きワールドアクションモデルマニピュレーション
- シミュレータ非依存の布操作のための簡易グリッパインタフェースマニピュレーション
- WRAP: 治具不要の力覚考慮型マルチロボット組立計画マニピュレーション
- 受動的実行から能動的探索へ:実環境におけるエージェント型身体性マニピュレーションマニピュレーション
- 全手把持のためのリアルタイム力制御フレームワークマニピュレーション
- 剛体-空気圧ハイブリッドマニピュレータの連成状態空間モデリング・制御・方策蒸留マニピュレーション