日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.02788

Skill2Real:ゼロショットSim-to-Realロボットマニピュレーションのためのエージェント型スキル学習

Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

提案・検証・統治ループでシミュレーション内のスキルを共通API上で学習し、実機へゼロショット転移するフレームワークを提案。LIBERO-90で学習したスキルが実世界の4タスクで78.75%の成功率を達成。

詳しい要約

1. どんなもの?

- シミュレーションで学習したロボット操作スキルを実機へゼロショット転移する枠組み Skill2Real を提案する。 - 共通の application programming interface (API) を通じて実行可能なスキルを学習する agentic policy framework。 - Proposer-Verifier-Governor (PVG) ループで privileged simulation evidence を使い結果を診断・更新検証する。 - Cerebellum が局所操作スキルを獲得し、Brain が Cerebellum を凍結してタスクレベル構成を学習する階層構造。 - 両メモリは task-policy fine-tuning や skill-memory updates なしで実機に転移する。

2. 先行研究と比べてどこがすごい?

- 従来の sim-to-real は知覚・動力学・embodiment の差に弱く、タスク知識の再学習が必要だった。 - 本手法は共通 API と公開観測に接地したスキルを学習し、実機での fine-tuning なしの転移を可能にする。 - LIBERO-90 で学習した凍結チェックポイントを GPT-6 Astra で評価すると、LIBERO-Pro Long 成功が 2.0% から 56.3% へ向上し、Pro Long で訓練していない。 - Robosuite でも Sol 85.1%、Opus 5 89.4% の平均成功率を7タスクで達成。 - 実機4タスクで凍結 Sol 学習 LIBERO-90 スキルが平均 78.75% 完了。

3. 技術・手法の肝は?

- Proposer-Verifier-Governor (PVG) ループが privileged simulation evidence で結果を診断し更新を検証する。 - 学習スキルを public observations と API semantics に接地させる。 - Cerebellum が局所操作スキルを獲得し、Brain が Cerebellum 凍結下でタスクレベル構成を学習する階層。 - 共通の application programming interface (API) を通じて実行可能スキルを表現する。 - 実機転移時に task-policy fine-tuning や skill-memory updates を行わない。

4. どうやって有効だと検証した?

- LIBERO-90 で GPT-5.6 Sol がスキル学習し、各凍結チェックポイントを GPT-6 Astra で評価。 - LIBERO-Pro Long 成功が 2.0% から 56.3% へ向上(Pro Long 未訓練)。 - Robosuite で Sol 85.1%、Opus 5 89.4% の平均成功率を7タスクで確認。 - 実機4操作タスクで凍結 Sol 学習 LIBERO-90 スキルが平均 78.75% 完了。 - Verifier または Governor を除くと LIBERO-90 訓練後の Pro Long 成功がそれぞれ 17.3、13.3 ポイント低下。

5. 議論はある?

- Verifier と Governor の寄与が ablation で示され、PVG ループの重要性が示唆される。 - 共通 API を通じた実行可能スキル階層の学習と転移が支持される。 - ただし実機評価は4タスクに限られ、より多様な embodiment や知覚差への一般性は要旨からは不明。 - 実機での fine-tuning なし転移の限界や失敗モードは要旨からは不明。

6. 次に読むべき論文は?

- LIBERO-90、LIBERO-Pro Long、Robosuite に関する研究。 - Proposer-Verifier-Governor (PVG) ループや privileged simulation evidence を用いる sim-to-real 手法。 - Cerebellum と Brain の階層的スキル学習・構成に関する研究。 - GPT-5.6 Sol、GPT-6 Astra、Opus 5 などの agentic policy 評価に関連する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xincheng He, Siyu Ma, Chang Yu, Yunuo Chen, Yanjia Huang, Ying Nian Wu, Yin Yang, Chenfanfu Jiang

分類: cs.RO

原文アブストラクト

Transferring robotic skills from simulation to reality requires task knowledge that remains usable across differences in perception, dynamics, and embodiment. We introduce Skill2Real, an agentic policy framework that learns executable skills through a shared application programming interface (API). A Proposer-Verifier-Governor (PVG) loop uses privileged simulation evidence to diagnose outcomes and validate updates, while keeping learned skills grounded in public observations and API semantics. The Cerebellum first acquires local manipulation skills; the Brain then learns task-level composition with the Cerebellum frozen. Both memories transfer to the real robot without task-policy fine-tuning or skill-memory updates. As GPT-5.6 Sol learns skills on LIBERO-90, evaluating each frozen checkpoint with GPT-6 Astra raises LIBERO-Pro Long success from 2.0% to 56.3%, without training on Pro Long. Independent Robosuite training reaches 85.1% and 89.4% mean success with Sol and Opus 5 across seven tasks, respectively. Frozen Sol-trained LIBERO-90 skills achieve 78.75% mean completion across four real-world manipulation tasks with Astra. Removing the Verifier or Governor during LIBERO-90 training lowers final Pro Long success by 17.3 and 13.3 percentage points, respectively. These results support learning and transferring a hierarchy of executable skills through a common robot interface.

関連論文

PR本紙発行元 EmplifAI