日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.02717

RoboBridge:シミュレーションから実世界への転移のための自己進化型具現化エージェントフレームワーク

RoboBridge: A Self-Evolving Embodied Agent Framework for Sim-to-Real Transfer

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーションと実世界で共有されるタスク知識を手続きとして表現し、対話フィードバックでスキルを自己進化させることで、VLAポリシーを再学習せずにsim-to-real転移を実現するフレームワークを提案。

詳しい要約

1. どんなもの?

- 本論文は、embodied intelligence の sim-to-real transfer を executable task skills の継続的適応として扱うフレームワーク RoboBridge を提案する。 - タスク知識を task intent、observations、tool operations、outcome verification を結ぶ procedures として表現する。 - 事前学習済み vision-language-action policy を再利用可能な action tool として公開し、推論時 guidance で強化する。 - シミュレーションと現実で共有される task semantics と interaction interfaces に転移可能スキルを接地する。 - LIBERO-PRO と対応する物理タスクで評価し、skill evolution と post-transfer adaptation を調べる。

2. 先行研究と比べてどこがすごい?

- 従来の end-to-end vision-language-action policies は強力な manipulation 能力を持つが、物理環境への転移には視覚・動力学条件の調整、追加の target-domain demonstrations 収集、追加訓練が必要だった。 - 既存の tool-using embodied agents はタスク実行と経験再利用を重視するが、シミュレーションと現実をまたぐ procedural knowledge の転移と継続的適応のサポートは限定的だった。 - RoboBridge は sim-to-real transfer を executable task skills の継続的適応として捉え、追加訓練なしで fine-grained execution を可能にする点が新しい。 - 転移可能スキルを task semantics と interaction interfaces に接地し、環境依存操作を実世界実行フィードバックで選択的に修正できる。

3. 技術・手法の肝は?

- タスク知識を procedures として表現し、task intent、observations、tool operations、outcome verification を接続する。 - 相互作用フィードバックを用いて候補 skill revisions を生成し、評価後に persist または reject する。 - 事前学習済み vision-language-action policy を再利用可能な action tool として公開し、推論時 guidance で強化する。 - これにより基盤 policy を再訓練せずに fine-grained execution を実現する。 - シミュレーションと現実で共有される task semantics と interaction interfaces に転移可能スキルを接地する。

4. どうやって有効だと検証した?

- LIBERO-PRO と対応する物理タスクでフレームワークを評価した。 - skill evolution と post-transfer adaptation の両方を調べた。 - 評価により、one-shot policy deployment から環境をまたぐ continual procedural learning への経路を提供することを示した。 - 具体的な定量的結果や比較指標は要旨からは不明。

5. 議論はある?

- 本フレームワークは、再利用可能なタスク構造を保持しつつ、環境依存操作を実世界実行フィードバックを通じて選択的に修正できる。 - これにより one-shot policy deployment から継続的な procedural learning への道筋を示す。 - ただし、skill revision の評価基準や失敗時の挙動、スケーラビリティ、他のタスクへの一般化可能性についての議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: end-to-end vision-language-action policies、tool-using embodied agents。 - 関連手法: vision-language-action policy、sim-to-real transfer、LIBERO-PRO。 - 同分野の定番として、domain randomization、system identification、imitation learning、reinforcement learning などが次に読むべき候補として挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chenxi Li, Zhangrui Zhao, Rui Li, Yuan Gao, Kehui Liu, Jiarui Li, Dong Wang, Tong Si, Minting Pan, Wanli Ouyang, Dongzhan Zhou

分類: cs.RO

原文アブストラクト

A key challenge in bringing embodied intelligence into the real world is transferring capabilities from simulation to reality and enabling agents to continually adapt after deployment. End-to-end vision-language-action policies provide strong manipulation capabilities, but their transfer to physical environments typically relies on calibrating simulated visual and dynamical conditions, collecting additional target-domain demonstrations, and optimizing the policy through further training. Tool-using embodied agents offer flexible task orchestration, yet existing systems primarily emphasize task execution and experience reuse within a given environment, with limited support for transferring procedural knowledge and continuously adapting it across simulation and reality. We propose RoboBridge, a framework that treats sim-to-real transfer as the continued adaptation of executable task skills. The agent represents task knowledge as procedures connecting task intent, observations, tool operations, and outcome verification. Interaction feedback is used to generate candidate skill revisions, which are evaluated before being persisted or rejected. A pretrained vision-language-action policy is exposed as a reusable action tool and enhanced with inference-time guidance, enabling fine-grained execution without retraining the underlying policy. RoboBridge grounds transferable skills in task semantics and interaction interfaces shared across simulation and reality. This representation preserves reusable task structure while allowing environment-dependent operations to be selectively revised through real-world execution feedback. We evaluate the framework on LIBERO-PRO and corresponding physical tasks, studying both skill evolution and post-transfer adaptation. Our framework provides a route from one-shot policy deployment to continual procedural learning across environments.

関連論文

PR本紙発行元 EmplifAI