日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.19661

ReShoot: 記録済みロボット実演の生成的視覚ドメインランダム化による視覚運動ポリシー学習

ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning

シェア:XThreadsFacebookLINEはてブBluesky

記録済みのロボット実演を視覚的に再レンダリングして外観を多様化し、追加データ収集なしで視覚運動ポリシーの頑健性を高める手法を提案。

詳しい要約

1. どんなもの?

- 記録済みのロボット実演を再レンダリングし、視覚的多様性を合成するフレームワーク ReShoot を提案。 - vision-language model がシーンを説明し、対象属性(背景・物体色・材質)を編集、edge-conditioned video generator が両カメラ視点を再レンダリング。 - 指示文も更新し、action sequence と proprioceptive trajectory は再ラベリングせずそのままコピー。 - 外観変化に対する visuomotor policy の過適合を、データ収集ではなく生成で緩和する。

2. 先行研究と比べてどこがすごい?

- 従来は新たな視覚条件ごとに追加実演を収集しており、ロボット・制御環境・人手を繰り返し要する点が高コスト。 - ReShoot は既存の記録済み実演を再レンダリングするため、負担を data collection から generation へ移す。 - 各生成エピソードは記録済みの action と proprioceptive ラベルを保持する点が特徴。 - 要旨からは不明: 他の visual domain randomization 手法との直接比較。

3. 技術・手法の肝は?

- vision-language model がシーンを caption し、背景・物体色・材質などの targeted attribute を編集。 - edge-conditioned video generator が編集内容に合わせて両カメラ視点を再レンダリング。 - 指示文を編集内容に応じて更新。 - action sequence と proprioceptive trajectory は verbatim にコピーし、再ラベリングしない。

4. どうやって有効だと検証した?

- LIBERO で、記録済み実演と再レンダリング実演を等量混合して学習した policy が、記録のみの学習と同等の性能(96.5% vs. 96.9%)。 - 混合学習セットは LIBERO-Plus での scene perturbations に対する robustness を改善(85.5% vs. 82.3%)。 - 2つの physical robotic platforms で、43件および100件の事前収集実演を用いた ReShoot の適用により、recolored objects での成功率が 0.0% から 42.9% および 47.5% に向上。 - 元の記録外観での性能は維持。

5. 議論はある?

- 要旨からは不明: 生成の計算コスト、生成品質の限界、失敗事例、他の外観変化への一般化。 - 要旨からは不明: 実機実験の詳細な設定や被験者数など。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: なし。 - 関連手法として、visual domain randomization、imitation learning、vision-language model、edge-conditioned video generation、LIBERO、LIBERO-Plus が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chiyoung Kim, Min Sung Choi, Jinho Ju, Chanhoe Gu, Donghwan Hwang, Wonseok Choi, Woongsun Jeon, Minhyeok Lee

分類: cs.RO

原文アブストラクト

Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each novel visual context; however, this approach is resource-intensive, requiring repeated access to a robot, a controlled environment, and human operation for every appearance condition to be covered. We introduce ReShoot, a framework that synthesizes visual diversity by re-rendering previously recorded demonstrations under altered appearances, thereby shifting the burden from data collection to generation. A vision-language model captions the scene, edits a targeted attribute (e.g., background, object color, or material), and an edge-conditioned video generator re-renders both camera views to match. The instruction is updated accordingly. The action sequence and proprioceptive trajectory are copied verbatim without relabeling, so each generated episode retains the recorded action and proprioceptive labels. On LIBERO, a policy trained on an equal mixture of recorded and re-rendered demonstrations matches the performance of recorded-only training (96.5% vs. 96.9%). Moreover, the mixed training set improves robustness to scene perturbations on LIBERO-Plus (85.5% vs. 82.3%). Across two physical robotic platforms, deploying ReShoot with 43 and 100 pre-collected demonstrations increased the success rate on recolored objects from 0.0% to 42.9% and 47.5%, respectively, while maintaining performance under the original recorded appearance.

関連論文

PR本紙発行元 EmplifAI