日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.27308

EmbodiedSWE:長期的器用ロボティクスのためのコーディングエージェント

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

シェア:XThreadsFacebookLINEはてブBluesky

コーディングエージェントが長期的・器用なロボットタスクを解き、その解を多様な軌道に拡張してVLAを訓練し、実機タスクまで完了できることを示した研究。

詳しい要約

1. どんなもの?

- 長期的・器用なロボティクス作業のためのcoding agentsを研究。 - その解決策が汎用ロボットポリシー学習のスケーラブルなsupervisionになり得るか問う。 - EMBODIEDSWE-BENCHを開発:接触豊富なmanipulation、deformable objects、最大30分の連続interactionを要するlong-horizonタスクを含むsimulation benchmark。 - EMBODIEDSWE-GENを導入:coding agentの単一解を多様なtrajectoriesに拡張しVLAを訓練。

2. 先行研究と比べてどこがすごい?

- 先行研究との具体的比較は要旨からは不明。 - frontier coding agentsが複雑なlong-horizonタスクを解き、タスクとembodimentを跨いでprior solutionsを転移できることを発見。 - しかし得られた解は反復的interactionを要し、個々のタスクインスタンスに特化しがち。 - EMBODIEDSWE-GENにより単一解から多様なtrajectoriesを生成し、VLA性能がdemonstrations数と共に向上。 - agent-aided diversificationがheld-outタスク変種へのgeneralizationを改善。

3. 技術・手法の肝は?

- EMBODIEDSWE-BENCH:接触豊富なmanipulation、deformable objects、long-horizonタスク(最大30分)を網羅するsimulation benchmark。 - coding agentsがこれらのタスクを解くための支援ツールを設計。 - EMBODIEDSWE-GEN:coding agentの単一解を大規模で多様なtrajectoriesに拡張し、VLAを訓練。 - agent-aided diversificationを実施。

4. どうやって有効だと検証した?

- frontier coding agentsがEMBODIEDSWE-BENCHの複雑なlong-horizonタスクを解けることを確認。 - タスクとembodimentを跨いだprior solutionsの転移を確認。 - VLA性能が生成demonstrations数の増加に伴い向上。 - agent-aided diversificationがheld-outタスク変種へのgeneralizationを改善。 - coding-agent-generated simulation demonstrationsのみでfinetuneしたVLAがreal robotでlong-horizonタスクを完了。

5. 議論はある?

- 得られた解は実用的だが、substantial iterative interactionを要し、個々のタスクインスタンスに特化する傾向。 - スケーラブルなsupervisionとしての可能性と限界が議論の焦点。 - 具体的な議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、VLA (Vision-Language-Action) models、coding agents、long-horizon robotics benchmarks、deformable object manipulation、contact-rich manipulationに関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoxiang You, Zeyu Shen, Yilang Liu, Zhicheng Zheng, Lihan Zha, Kashu Yamazaki, Mingtong Zhang, Suning Huang, Jiankai Sun, Qianzhong Chen, Lucy He, Kaiyuan Liu, Haoran Chang, Katerina Fragkiadaki, Dhruv Shah, Mac Schwager, Peter Henderson, Ian Abraham, Canwen Xu

分類: cs.RO

原文アブストラクト

We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both tasks and embodiments. We also design supporting tools that help agents more effectively solve these tasks. However, the resulting solutions require substantial iterative interaction and are typically specialized to individual task instances. We therefore introduce EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA. VLA performance improves with more generated demonstrations, and agent-aided diversification improves generalization to held-out task variations. We also show that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot. Together, our framework uses coding agents to solve complex robotics tasks and turn verified solutions into scalable supervision for robot policies.

関連論文

PR本紙発行元 EmplifAI