日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04929

RobotUse: 計算・文脈・意思決定の割り当てによるロボットエージェント

RobotUse: Allocating Computation, Context, and Decisions

シェア:XThreadsFacebookLINEはてブBluesky

ロボットエージェントが視覚的に目標と姿勢を選び、バックエンドが幾何・運動計画・制御を担うハーネスを提案。サブエージェントが詳細を保持し、永続的なプレイブックで実行から学習する。

詳しい要約

1. どんなもの?

- ロボットエージェントが意図した行動と観測結果を結びつけ、反復試行で選択を修正するための文脈を保持することを支援する - 計算・文脈・意思決定を物理行動の指定と修正の周りに組織化する robot agent harness『RobotUse』 - agent は視覚的に target と pose を選択し、backend が geometry・motion planning・control を担当 - subagent が各 subgoal 内の詳細な interaction を保持し、後続の意思決定に必要な情報を返す - continual harnessing により persistent playbook を更新し、実行から学習する

2. 先行研究と比べてどこがすごい?

- 既存 interface は選択肢を predefined tool 内に閉じ込めたり、agent に詳細な実行 code と増大する履歴の管理を要求していた - RobotUse は computation・context・decisions を物理行動の指定と修正の周りに組織化する点が異なる - RoboLab で 45% の task success を達成し、CaP-X を 6.7 percentage points 上回る - コンパクトな decision context を維持し、predefined action abstraction への依存を減らす - 不完全な feedback 下でも real-world execution から学習し、後続タスクへ転移できることを示す

3. 技術・手法の肝は?

- robot agent harness として computation・context・decisions を物理行動の指定と修正の周りに組織化 - agent が視覚的に target と pose を選択し、backend が geometry・motion planning・control を処理 - subagent が各 subgoal 内の詳細な interaction を保持し、後続判断に必要な情報のみを返す - continual harnessing により persistent playbook を更新し、実行経験から学習 - これにより decision context をコンパクトに保ちつつ、predefined action abstraction への依存を低減

4. どうやって有効だと検証した?

- RoboLab 上で評価し、45% の task success を達成 - CaP-X を 6.7 percentage points 上回る性能を確認 - コンパクトな decision context の維持と predefined action abstraction への依存低減を確認 - 不完全な feedback を含む real-world execution から学習できることを示す - 学習した内容を後続タスクへ転移できることを示す

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- CaP-X(要旨で比較対象として言及) - RoboLab(要旨で評価に用いられたと記載) - robot agent harness や predefined action abstraction に関する関連研究(要旨に具体的な参照は無いため、同分野の一般名として挙げる)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junhoo Lee, Injun Baek, Seungyeon Kim, Suhyun Jeon, Minkyu Kim, Baekseung Kim, Nojun Kwak

分類: cs.RO

原文アブストラクト

Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.

関連論文

PR本紙発行元 EmplifAI