日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12386

ARC: ロボット基盤モデルのための推論レシピ

ARC: A Reasoning Recipe for Robot Foundation Models

シェア:XThreadsFacebookLINEはてブBluesky

既存のロボット基盤モデルに「推論トレース」を自動生成して学習させることで、追加のロボット実演データや大規模学習なしにゼロショット性能を大幅に向上させる手法ARCを提案。

詳しい要約

1. どんなもの?

ロボット基盤モデル(RFM)のzero-shotタスク性能を、大規模モデルや追加のロボットデモンストレーション、大規模学習に頼らず改善する推論レシピARCを提案する。ARCは3要素からなる。 - reasoning trace - scalable automatic labeling pipeline - pretrained RFMをtraceで制御するための適応戦略 - 既存デモンストレーションから自動生成し、ARC-Trace-DROIDをDROIDから構築 - VLAs($π_{0.5}$)やWAMs(Cosmos3-Nano-Policy)に適用可能

2. 先行研究と比べてどこがすごい?

従来はより大きなモデル、より多くのロボットデモ、高コストな大規模学習に依存していた。 - ARCは補完的アプローチとして、適切なreasoning recipeで既存SOTA RFMのzero-shot性能を大幅改善 - 追加のロボットデモンストレーションやfoundation-scale trainingなしで、これまでにない性能向上を達成 - RoboLab-120とMolmoSpacesで新たなSOTAを確立 - RoboLab-Reasoning-50で最大50 percentage points向上 - 実機で$π_{0.5}$のタスク成功率を82.2 percentage points改善

3. 技術・手法の肝は?

ARCの肝は3つの要素。 - reasoning trace: ロボットのnext actionに基づき、その因果構造(なぜ適切か、どんな効果を生むか)を説明する - scalable automatic labeling pipeline: 既存デモンストレーションからtraceを自動生成し、新規ロボットデータ収集を不要にする(ARC-Trace-DROIDをDROIDから構築) - 適応戦略: SOTA VLAs($π_{0.5}$)やWAMs(Cosmos3-Nano-Policy)がtraceを制御に使えるよう、各モデルのアーキテクチャと能力に合わせたfine-tuningとinferenceを実施

4. どうやって有効だと検証した?

RoboLab-120とMolmoSpacesで評価し、新たなSOTAを確立。 - RoboLab-Reasoning-50で最大50 percentage pointsの向上 - 実ロボットで$π_{0.5}$のタスク成功率が82.2 percentage points改善 - 追加のロボットデモンストレーションやfoundation-scale trainingなしでzero-shot性能向上を確認

5. 議論はある?

要旨からは不明。 - 限界や失敗事例、計算コスト、汎化性に関する議論は明記されていない - 有効性の主張と検証結果が中心

6. 次に読むべき論文は?

要旨で参照・比較されている研究や関連手法。 - $π_{0.5}$(VLA) - Cosmos3-Nano-Policy(WAM) - DROID(データセット) - RoboLab-120、RoboLab-Reasoning-50、MolmoSpaces(ベンチマーク) - これらに加え、同分野の定番としてRT-2、OpenVLA、Octoなどが次に読む候補

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Gokul Puthumanaillam, Tao Sun, Elie Aljalbout, Moritz Reuss, Zhaoshuo Li, Fabio Ramos, Ankit Goyal, Jenai Xuning Yang

分類: cs.RO, cs.AI

原文アブストラクト

The prevailing approach to improving robot foundation models (RFMs) relies on larger models, more robot demonstrations, and costly training at scale. We show that there exists an effective and efficient complementary approach: the right reasoning recipe can substantially improve the zero-shot task performance of existing state-of-the-art RFMs. We refer to this recipe as ARC. It consists of three key ingredients: a reasoning trace, a scalable automatic labeling pipeline, and a strategy for adapting pretrained RFMs to use these traces for control. First, we find that effective reasoning traces should be grounded in the robot's next action and explain its causal structure: why the action is appropriate and what effect it should produce. Second, we show that these traces can be generated automatically from existing demonstrations, enabling us to construct ARC-Trace-DROID from DROID without collecting new robot data. Third, we show how state-of-the-art VLAs such as $π_{0.5}$ and WAMs such as Cosmos3-Nano-Policy can learn to use these traces for control, with fine-tuning and inference tailored to each model's architecture and capabilities. Using ARC, we obtain gains in zero-shot RFM performance that, to our knowledge, are unprecedented without additional robot demonstrations or foundation-scale training. The adapted models establish a new state of the art on RoboLab-120 and MolmoSpaces, with gains of up to 50 percentage points on RoboLab-Reasoning-50. On real robots, ARC improves $π_{0.5}$'s task success by 82.2 percentage points. Project website: https://arc-robot-reasoning.github.io/

関連論文

PR本紙発行元 EmplifAI