日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.01939

トークンを減らして行動を改善:成功率14%向上・トークン65%削減を実現するGPT-6 Astraロボットエージェント

Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

シェア:XThreadsFacebookLINEはてブBluesky

ロボットプリミティブとVLAポリシーをPythonコードとして組み合わせ、必要な観測だけを要求するフレームワークPyRUA-Leanを提案。同一プランナーで成功率を63.1%から71.7%に向上させつつ、入力トークンを65%削減した。

詳しい要約

1. どんなもの?

- 視覚言語モデル(VLM)エージェントによるロボット制御のトークンオーバーヘッドを削減するフレームワーク「PyRUA-Lean」を提案。 - フィードバック駆動型のプリミティブ合成と選択的観測を組み合わせ、Pythonセル内で条件分岐やローカルリトライを実行。 - 明示的に要求された画像と状態フィードバックのみを返し、再計画に利用。 - LIBERO-PRO、RoboTwin 2.0、RoboCasa365の700タスクで評価。

2. 先行研究と比べてどこがすごい?

- 同じGPT-6 Astraプランナーとロボットプリミティブを用いるツール呼び出しベースラインと比較。 - 同等のLLM呼び出し予算下で、全体成功率を63.1%から71.7%に向上。 - 両エージェントが解決したインスタンスでは、LLM呼び出しが49%減、入力トークンが65%減。

3. 技術・手法の肝は?

- フィードバック駆動型のプリミティブ合成と選択的観測を結合。 - エージェントが古典的ロボットプリミティブと学習済みVLAポリシーをPythonセルに合成。 - Pythonセル内で条件チェックとローカルリトライを実行。 - 再計画のために明示的に要求された画像と状態フィードバックのみを返す。

4. どうやって有効だと検証した?

- LIBERO-PRO、RoboTwin 2.0、RoboCasa365から700のシミュレーションタスクインスタンスで評価。 - 同一のGPT-6 Astraプランナーとロボットプリミティブを用いるツール呼び出しベースラインと比較。 - 成功率、LLM呼び出し数、入力トークン数を指標として検証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:ツール呼び出しベースライン(具体的名称は要旨に記載なし)。 - 関連手法:VLAポリシー、LIBERO-PRO、RoboTwin 2.0、RoboCasa365。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou

分類: cs.CV

原文アブストラクト

Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.

関連論文

PR本紙発行元 EmplifAI