日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.35432

自己進化するコーディングエージェント:デジタルプログラムから物理世界の知能へ

Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

シェア:XThreadsFacebookLINEはてブBluesky

タスクの状態と実行をコードとして表現する「Physical Coding」を提案し、知覚・計画・制御ツールを呼び出して外部フィードバックから自己修正するエージェントHexaAnythingを構築した。RoboCasa365でVLAを上回る性能を示した。

詳しい要約

1. どんなもの?

VLAやWAMモデルは観測と指示を直接ロボット行動に写像するが、レイアウトや視点の変化に弱く指示の汎化も乏しい。原因はタスク要件・条件・進捗・失敗回復が行動列に暗黙的に埋め込まれ、検査や修正が困難な点にある。本研究はPhysical Codingを提案し、タスク状態と実行をコードとして表現する。Code as Worldは物体・関係・制約・進捗を記録し、Code as Policyは計画・検証・回復・実行を組織する。HexaAnythingは知覚・計画・制御ツール(VLA/WAMポリシー含む)を呼び出し、外部フィードバックからループ内意思決定を行う。検証済みトレースはデータと記憶になり、ツールとHarnessからモデル重み・アーキテクチャ・最終的にハードウェアとタスク設計への進化を可能にする。

2. 先行研究と比べてどこがすごい?

VLA/WAMは観測と指示を直接行動に写像するため、ポリシーが訓練に縛られ、わずかなレイアウトや視点の変化で失敗し、指示の汎化も悪い。本研究は表現の根本原因に着目し、タスク状態と実行をコードとして明示化する。これにより検査・修正可能な手続き、明示的状態、管理可能な実行を実現し、物理世界での汎化と長期的実行を可能にする。デジタルコーディングエージェントの前例(LLMがツールを呼び、結果を検証し、フィードバックから改訂)を物理世界に持ち込む点が新しい。RoboCasa365でComposite-Unseenと全体成功率がXR-1 VLAを上回り、Harness-trained HexaModelは全splitでベースを上回る。

3. 技術・手法の肝は?

Physical Codingの中核は、タスク状態と実行をコードで表現すること。Code as Worldが物体・関係・制約・進捗を記録し、Code as Policyが計画・検証・回復・実行を組織する。HexaAnythingは知覚・計画・制御ツール(VLA/WAMポリシーを含む)を呼び出し、外部フィードバックからループ内で意思決定する。検証済みトレースはデータと記憶として蓄積され、ツールとHarnessからモデル重み、アーキテクチャ、最終的にハードウェアとタスク設計への進化を可能にする。

4. どうやって有効だと検証した?

RoboCasa365でHexaAnythingがComposite-Unseenと全体成功率でXR-1 VLAを上回り、Harness-trained HexaModelが全splitでベースを上回った。PhyBenchとdual-arm AgileXロボットで、エージェントが物理実験を自律完了し、ほとんどの卓上タスクを公開結果よりしばしば高速に遂行した。データ・モデル・ツールの自己進化が観察された。

5. 議論はある?

データ・モデル・ツールの自己進化が観察された。今後の課題は重みへの内部化、アーキテクチャ・言語・表現・タスクの自律的再設計、製造業と科学への展開。

6. 次に読むべき論文は?

XR-1 VLA、RoboCasa365、PhyBench、dual-arm AgileX robot、VLA、WAM、Harness、HexaModel。関連する同分野の定番としてRT-2、OpenVLA、Diffusion Policyなど。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hongcheng Gao, Jingjing Zhou, Zelin Zheng, Shijia Ge, Jay Zhu, Yazhe Wang, Jianshu Zeng, Xuan Shangguan, Di Wu, Lingyu He, Zhiqi Jia, Sihang Wu, Xiao He

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, progress, and failure recovery are implicitly encoded in action sequences, making them difficult to inspect or revise. Digital coding agents offer a precedent: LLMs call tools, verify results, and revise from feedback as executable code. The same working pattern of explicit state, manageable execution, and revisable procedures underlies generalization and long-horizon execution in the physical world, letting physical experience return as reusable programs, memory, or evidence. We propose Physical Coding, representing task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes planning, verification, recovery, and execution. We build HexaAnything, which calls perception, planning, and control tools, including VLA/WAM policies, and makes in-the-loop decisions from external feedback. Verified traces become data and memory, enabling evolution from tools and Harness to model weights, architectures, and ultimately hardware and task design. On RoboCasa365, HexaAnything improves Composite-Unseen and overall success over XR-1 VLA, and its Harness-trained HexaModel beats the base on every split, indicating code traces internalize physical execution. On PhyBench and a dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results. We observe data, model, and tool self-evolution; future work targets weight internalization, autonomous redesign of architectures, languages, representations, and tasks, and deployment in manufacturing and science.

関連論文

PR本紙発行元 EmplifAI