ReShoot: 記録済みロボット実演の生成的視覚ドメインランダム化による視覚運動ポリシー学習
ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning
記録済みのロボット実演を視覚的に再レンダリングして外観を多様化し、追加データ収集なしで視覚運動ポリシーの頑健性を高める手法を提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Chiyoung Kim, Min Sung Choi, Jinho Ju, Chanhoe Gu, Donghwan Hwang, Wonseok Choi, Woongsun Jeon, Minhyeok Lee
分類: cs.RO
原文アブストラクト
Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each novel visual context; however, this approach is resource-intensive, requiring repeated access to a robot, a controlled environment, and human operation for every appearance condition to be covered. We introduce ReShoot, a framework that synthesizes visual diversity by re-rendering previously recorded demonstrations under altered appearances, thereby shifting the burden from data collection to generation. A vision-language model captions the scene, edits a targeted attribute (e.g., background, object color, or material), and an edge-conditioned video generator re-renders both camera views to match. The instruction is updated accordingly. The action sequence and proprioceptive trajectory are copied verbatim without relabeling, so each generated episode retains the recorded action and proprioceptive labels. On LIBERO, a policy trained on an equal mixture of recorded and re-rendered demonstrations matches the performance of recorded-only training (96.5% vs. 96.9%). Moreover, the mixed training set improves robustness to scene perturbations on LIBERO-Plus (85.5% vs. 82.3%). Across two physical robotic platforms, deploying ReShoot with 43 and 100 pre-collected demonstrations increased the success rate on recolored objects from 0.0% to 42.9% and 47.5%, respectively, while maintaining performance under the original recorded appearance.
関連論文
- チャンク型VLAマニピュレーションポリシーの学習と実機展開のためのSim-to-Real統合パイプラインsim2real
- 運動学を超えて:筋駆動模倣学習のためのシミュレーション忠実度ベンチマークsim2real
- CRISP: 多様な形状と接触ソルバを備えた接触リッチロボットシミュレーション基盤sim2real
- 同じ世界、異なる知識:孤立評価が世界モデルの修復を誤判定するときsim2real
- DEXTERA: 単一画像から実機展開可能な巧みなマニピュレーションへ向けたReal-to-Sim-to-Realsim2real
- 単一スキャンからのガウシアンスプラッティングによる実演合成と視覚運動ポリシー学習sim2real