日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11480

RoboAware: 反事実的結果から身体スキルの協調を学習する

RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

シェア:XThreadsFacebookLINEはてブBluesky

コーディングエージェントが複数のロボットスキルを組み合わせる際に、状態に応じてどのポリシー系統が成功するかを反事実的結果から学習する調整器を提案し、100タスクで77%の成功率を達成した。

詳しい要約

1. どんなもの?

- ロボットの modular skills と frozen end-to-end policies を coding agent が組み合わせる枠組み - 現在の physical state でどの policy family が成功するかを予測する責任調整器を学習 - 状態条件付き responsibility coordinator のみを counterfactual outcomes から学習 - REPL に着想を得た P^5 schema と階層的 MDP を提案 - 展開時は観測文脈に応じて policy family を選び、frozen coding agent が次の local code block を生成

2. 先行研究と比べてどこがすごい?

- 既存の code-as-policy や VLA-harness ベースラインを上回る - 既存研究に欠けていた counterfactual branch outcomes を導入 - 選択された分岐の経験では観測されない結果を露出させる SCB を提案 - 単一エピソード評価で全体成功率 77.0%、RoboSuite 90.0%、LIBERO-Pro 73.8%、RoboTwin 90.0% を達成 - 従来は coding agent の skill orchestration に留まっていたが、状態条件付き調整を学習

3. 技術・手法の肝は?

- P^5 schema で skills を 5 つの semantic stages に均一に整理し、責任比較の場を定義 - 階層的 MDP を P^5 に基づき定式化 - State-Locked Counterfactual Branching (SCB): 同じ training state を復元し、各 admissible family から code block を生成・実行 - Execution-Aware Learning (EAL): Monte Carlo tree search と Q-learning を組み合わせ、結果を family-conditioned values に蒸留 - 展開時は coordinator が観測可能な context に応じて policy family を選択

4. どうやって有効だと検証した?

- 100 タスクの単一エピソード評価を包括的に実施 - 全体成功率 77.0% - RoboSuite で SOTA 平均 90.0% - 多様な LIBERO-Pro task clusters で 73.8% - 難易度の高い RoboTwin bimanual tasks で 90.0% - 既存の code-as-policy および VLA-harness ベースラインと比較して優位

5. 議論はある?

- 要旨からは不明 - 限界や失敗事例、計算コスト、スケーラビリティに関する議論は要旨に記載なし

6. 次に読むべき論文は?

- REPL - code-as-policy - VLA-harness - Monte Carlo tree search - Q-learning - RoboSuite - LIBERO-Pro - RoboTwin

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bohan Zhou, Xingbei Chen, Emily Huang, Weilin Ruan, Haojian Huang, Yehang Zhang, Zexi Li, Wenqian Li, Qize Yu, Zetian Song, Leyi Wu, Jinghao Li, Mingxuan Song, Xinrun Xu, Zongyang Qiu, Yangkai Wei, Tianyi Zhang, Kaiwen Zhou, Yinchuan Li, James Cheng

分類: cs.RO, cs.AI

原文アブストラクト

Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the $P^5$ schema and formulate a hierarchical MDP based on it. $P^5$ organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.

関連論文

PR本紙発行元 EmplifAI