日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.37089

Real2Gym: 動画からロボット用トレーニング環境を構築し、実機へスキルを転移

Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

シェア:XThreadsFacebookLINEはてブBluesky

人間やロボットのデモ動画から物理的に整合するシミュレーション環境を再構築し、その中で操作スキルを学習・蒸留して実機ロボットへ転移するReal2Sim2Realフレームワークを提案。

詳しい要約

1. どんなもの?

- 人間やロボットの実演動画から、操作スキルを学習・再利用できる対話型シミュレーション環境(gym)を構築し、そのスキルを実機ロボットに転移する枠組み。 - Real2Sim2Real をエージェント的に統合し、Real2Sim モジュールで編集可能なシーン再構成、物体・カメラの位置合わせ、実演/リターゲット行動の物理検証、タスク条件付きバリエーション生成を行う。 - シミュレーション内でエージェントが操作段階の実行コードを生成し、結果を観察して成功・失敗を再利用可能なタスク手順、物体相対運動、回復戦略に蒸留する。 - 共有の知覚・制御インターフェースを通じて、モデル重みを更新せずに現在の観測に適応した運動をシミュレーションと実機で案内する。

2. 先行研究と比べてどこがすごい?

- 従来の Real2Sim や Sim2Real の個別手法と異なり、動画から対話型 gym を構築し、スキル獲得と実機転移までを一貫して行うエージェント的枠組みを提案。 - シミュレーション環境再構成の高忠実性を実現し、GPT-6 Astra Direct Mode と比較して成功率で 16.7% 向上、ポリシー実行トークン数を約 74.9% 削減。 - 実機 Franka ロボットでの 4 タスクにおいて、物理実行成功率で GPT-6 Astra Direct Mode を 33.3% 上回る。

3. 技術・手法の肝は?

- Real2Sim モジュール:編集可能なシーン再構成、入力との物体・カメラ位置合わせ、実演またはリターゲット行動のネイティブ物理実行による検証、行動実現可能性チェック付きのタスク条件付きバリエーション生成。 - エージェントによる操作段階の実行コード生成と結果観察、成功・失敗の蒸留による再利用可能なタスク手順、物体相対運動、回復戦略の獲得。 - 共有の知覚・制御インターフェースを通じ、モデル重みを更新せずに現在の観測に運動を適応させ、シミュレーションと実機でスキルを案内。

4. どうやって有効だと検証した?

- 広範な評価により、Real2Gym が高忠実度のシミュレーション環境再構成を可能にすることを実証。 - GPT-6 Astra Direct Mode と比較して、シミュレーション環境での成功率が 16.7% 高く、ポリシー実行トークン数が約 74.9% 少ない。 - 実機 Franka ロボットでの 4 タスクにおいて、物理実行成功率が GPT-6 Astra Direct Mode を 33.3% 上回る。

5. 議論はある?

- 要旨からは不明。 - 限界や今後の課題、倫理的影響、計算コスト、スケーラビリティなどに関する議論は要旨に記載されていない。

6. 次に読むべき論文は?

- GPT-6 Astra Direct Mode(比較対象として言及) - Real2Sim、Sim2Real に関する関連研究(具体的な論文名は要旨に記載なし) - 同分野の定番として、Domain Randomization、Imitation Learning、Reinforcement Learning などの一般手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu, Bo Dai, Mulin Yu, Tao Lu

分類: cs.CV

原文アブストラクト

Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.

関連論文

PR本紙発行元 EmplifAI