日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12369

身体化チューリングマシン:ロボットの再帰的自己改善のための状態を持つコード

Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの状態をコードで明示的に表現し、VLMやVLAを使わずコードのみで意思決定を行う手法COAPを提案。42の両腕タスクで70.24%の成功率を達成した。

詳しい要約

1. どんなもの?

- 提案:Embodied Turing Machine の視点で、ロボット政策をコードのみで記述する Code-Only-as-Policy (COAP) を提案。 - 特徴:VLA や VLM をループ内に持たず、カメラ画像と固有受容感覚から状態をコードで計測・追跡し、全決定をコードで行う。 - 適用:同一コードが複数エピソードに適用され、異なるタスクが VLM/VLA なしで一つのライブラリを共有。 - 目的:Recursive Self-Improvement (RSI) の媒体として、コーディングエージェントが閉ループでライブラリを開発。

2. 先行研究と比べてどこがすごい?

- 先行:VLA が観測を行動に写像し、Agent Harness (Agent-as-Policy, Harness VLA) が実行時に VLM を照会する。 - 利点(i) Explicit State:状態をコードに保存可能。 - 利点(ii) Execution:コードにより決定が制御可能、障害から柔軟に回復、オンラインで高速・低コスト。 - 利点(iii) Extensibility:新タスクが共有ライブラリを再利用・継承・拡張し、能力がタスク間で蓄積。 - 結果:RoboDojo の42両腕タスクで、テスト時にモデルなしで成功率70.24%を達成。

3. 技術・手法の肝は?

- 核心:Embodied Turing Machine のテープをロボットと環境の状態、ルールを政策と見なし、状態を正確に表現できれば決定を全てコードで記述。 - 実装:コードがカメラ画像と固有受容感覚からロボット・環境・タスク状態を計測・追跡し、それに基づき全決定を下す。 - 共有:同一コードがエピソード間で適用され、異なるタスクが VLM/VLA なしで一つのライブラリを共有。 - 自己改善:コーディングエージェントが閉ループでライブラリを開発し、各変更が明示的で制御可能。

4. どうやって有効だと検証した?

- データセット:RoboDojo の42両腕タスクで評価。 - 結果:テスト時にモデルなしで成功率70.24%を達成。 - 限界:COAP の上限は、決定のための状態表現の正確さとコードロジックの堅牢性に依存。 - 追加:エピソード間で適用可能なため、VLA や Agent Harness の効率的なデータエンジンにもなり得る。

5. 議論はある?

- 利点:Explicit State、Execution、Extensibility の3点を分析。 - 限界:状態表現の正確さとコードロジックの堅牢性が上限を決める。 - 位置づけ:COAP を embodied タスクの新パラダイムとして提案。 - 応用:VLA や Agent Harness の効率的なデータエンジンとしても機能し得る。

6. 次に読むべき論文は?

- 比較対象:VLA、Agent Harness (Agent-as-Policy, Harness VLA)、VLM。 - 関連手法:Recursive Self-Improvement (RSI)、Embodied Turing Machine。 - データセット:RoboDojo。 - 同分野の定番:Vision-Language-Action (VLA) モデル、Vision-Language Model (VLM) を用いたロボット政策。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kairui Hu, Siyuan Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu

分類: cs.RO, cs.CV

原文アブストラクト

Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.

関連論文

PR本紙発行元 EmplifAI