日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.09254

RoboRender: ロボット指向の動画生成による視覚的Sim-to-Real転移

RoboRender: Robot-Oriented Video Generation for Visual Sim-to-Real Transfer

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーションの深度動画・言語指示・ロボットマスクから写実的なRGB動画を生成し、実機へゼロショット転移可能な方策を学習するフレームワーク。

詳しい要約

1. どんなもの?

- シミュレーション軌跡をphotorealisticなRGB動画に変換するRoboRenderを提案。 - 条件はsimulated depth videos、language instructions、robot RGB mask videos。 - simulator geometry、robot motion、action labelsを保持しつつ、textures、backgrounds、distractorsを合成。 - 生成RGB動画とsimulator提供のstates/actionsを組にしてpolicyを学習し、zero-shot real-world deploymentを目指す。

2. 先行研究と比べてどこがすごい?

- 既存手法はintermediate representationsに依存し、rich semantic informationを失う、またはdeployment時に追加のperception modulesを要する。 - RoboRenderは中間表現を介さずphotorealistic RGB videosを直接生成し、policy学習に用いる。 - 実世界実験で、raw simulation renderings比約7.1x、conventional visual domain randomization比約3.6xの成功率向上。

3. 技術・手法の肝は?

- robot-oriented video generation modelを、simulated depth videos、language instructions、robot RGB mask videosで条件付け。 - simulator geometry、robot motion、action labelsを保持しながら、realistic textures、backgrounds、distractorsを合成。 - 生成RGB動画とsimulator-provided states/actionsをpairにしてpolicyを訓練。

4. どうやって有効だと検証した?

- robot video test setsで、depth-conditioned video generation baselinesより生成品質が優れる。 - 実世界でpick-and-place、articulated-object manipulation、mobile manipulationを検証。 - RoboRender生成データで訓練したpolicyが平均71%の成功率。 - 1軌跡あたりの生成動画数を増やすとpolicy性能が向上し、opening-task成功率が65 percentage points増加。

5. 議論はある?

- 生成video renderingがvisual sim-to-real gapを緩和し、zero-shot policy transferを可能にすることを示す。 - 生成動画数の増加が性能向上に寄与する点を報告。 - 限界や失敗事例、計算コスト、他タスクへの一般化については要旨からは不明。

6. 次に読むべき論文は?

- depth-conditioned video generation baselines(具体的名称は要旨からは不明)。 - conventional visual domain randomization。 - raw simulation renderings。 - 関連するvideo generationやsim-to-real transferの定番手法(例:video diffusion models、domain randomization)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Huang Huang, Wensi Ai, Ziyu Chen, Youhui Wang, Zijian Du, Yang Liu, Jiaolong Yang, Li Fei-Fei, Jiajun Wu

分類: cs.RO

原文アブストラクト

Simulation enables large-scale, low-cost robot data generation, but policies trained in simulation often fail to transfer to the real world due to the sim-to-real visual discrepancies. Existing approaches often rely on intermediate representations, which can discard rich semantic information or require additional perception modules at deployment. We address this visual sim-to-real gap with RoboRender, a framework that converts simulated trajectories into photorealistic RGB videos for policy learning. RoboRender trains a robot-oriented video generation model conditioned on simulated depth videos, language instructions, and robot RGB mask videos, preserving simulator geometry, robot motion, and action labels while synthesizing realistic textures, backgrounds, and distractors. The generated RGB videos are paired with simulator-provided states and actions to train policies for zero-shot real-world deployment. On robot video test sets, our video model outperforms depth-conditioned video generation baselines in generation quality. In real-world experiments across pick-and-place, articulated-object manipulation, and mobile manipulation tasks, policies trained on RoboRender-generated data achieve a 71% average success rate, outperforming raw simulation renderings and conventional visual domain randomization by approximately 7.1x and 3.6x, respectively. We further show that policy performance improves with more generated videos per simulation trajectory, increasing opening-task success by 65 percentage points. These results demonstrate that generative video rendering mitigates the visual sim-to-real gap for zero-shot policy transfer. Project website: https://robo-render.github.io/.

関連論文

PR本紙発行元 EmplifAI