日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.10479

Agentic RSR: シーン再構成と実行に基づくロボット政策によるReal-to-Sim-to-Real

Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

作業場の動画からシーンを再構成し、コーディングエージェントが実行可能な政策を開発、実機ロボットへ転移するフレームワークを提案。実機でのタスク成功率はシミュレーションの80%に達した。

詳しい要約

1. どんなもの?

- 実ロボット作業空間のビデオとタスク記述、既知のロボットモデルから、シーン再構成・ポリシー開発・実機実行を同一タスクで連結するフレームワーク。 - エージェントがメートルスケールを復元し、視覚フィードバックでシーンを反復改良し、MuJoCoでタスク関連相互作用を確認。 - コーディングエージェントが実行可能ポリシーを開発し、特権物体姿勢から視覚観測・ランダム化シミュレーションへ進展。 - ポリシーは1回の呼び出し内で複数観測と行動を交互に実行可能。エージェントは実行フィードバックで継続・再試行・修正。 - 共有タスクレベルインターフェースがポリシーと蓄積経験を実機に運び、新観測と安全チェックで実行を導く。

2. 先行研究と比べてどこがすごい?

- 従来はシーン再構成とポリシー開発が別々に扱われることが多かったが、本手法は同一操作タスクを通じて両者を連結。 - 再構成から実機実行までを一貫してエージェントが担い、実行フィードバックをポリシー改良に活用する点が特徴。 - ポリシーが1回の呼び出し内で複数観測・行動を交互に実行できる点も従来と異なる。 - 実機性能がシミュレーション性能を大幅に保持(成功率80%)することを示した。

3. 技術・手法の肝は?

- 作業空間ビデオ、タスク記述、既知ロボットモデルを入力。 - エージェントがメートルスケールを復元し、視覚フィードバックでシーンを反復改良。 - MuJoCoでタスク関連相互作用をチェック。 - コーディングエージェントが特権物体姿勢から視覚観測・ランダム化シミュレーションへと段階的にポリシーを開発。 - ポリシーは1回の呼び出し内で複数観測と行動を交互に実行可能。 - 実行フィードバックに基づきエージェントが継続・再試行・修正。 - 共有タスクレベルインターフェースでポリシーと経験を実機に転送し、新観測と安全チェックで実行。

4. どうやって有効だと検証した?

- 2種類のロボットを含む18の再構成シーンで評価。 - 4視点のDepth MAE平均0.1057 m、Lab ΔE76平均11.04、グレースケールSSIM平均0.6990。 - 実機実験で集計タスク成功率がシミュレーション成功率の80%に達し、シミュレーション性能の実機保持を示す。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- MuJoCo、Real-to-Sim-to-Real、Sim-to-Real転送、ビジョンに基づくロボットポリシー学習に関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yihan Li, Yating Feng, Shengjiu Sun, Jianing Chen, Hao Ren, Bowen Yang, Weisheng Xu, Qiwei Wu, Hui Cheng, Renjing Xu

分類: cs.RO, cs.CV

原文アブストラクト

A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab $ΔE_{76}$ is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.

関連論文

PR本紙発行元 EmplifAI