日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.02398

効果を保ち、実行者を捨てる:プログラム可能な効果-実行ワールドアクションモデル

Keep the Effect, Drop the Actor: Programmable Effect-to-Execution World-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

物体の動き(効果)だけをプログラム化し、それを任意のロボットで実行できるワールドアクションモデルPEWAMを提案。単一のデモから効果プログラムを生成し、4種のアームで再学習なしに動作する。

詳しい要約

1. どんなもの?

- ロボットデモンストレーションから、物体の動き(effect)とロボットの実行(execution)を分離し、effectのみを条件とする世界行動モデルPEWAMを提案。 - デモンストレーションをeffect program(物体の3Dキーポイント軌跡、把持点、終端状態)にコンパイルし、PEWAMがeffect、robot execution、action、terminal stateの4ストリームを独立したflow-matching時間で生成。 - プログラムを固定して実行をサンプリングすることで、推論をプログラミングとして扱い、ライブシーンから閉ループで再解決可能。

2. 先行研究と比べてどこがすごい?

- 従来のデモンストレーション条件付き手法や、ゴール画像・言語条件付きの同一バックボーンと比較して、LIBERO-Goalの未見タスクで1デモのプログラムが90エピソード中40を完了(最大19)。 - Meta-Worldの未見3クラスで公開最高結果を上回る(5クラス平均では及ばない)。 - 同じプログラムが再学習なしで4つのロボットアームで動作し、押し動作後の再解決で60エピソード中23を完了(デモ再生は6)。 - Frankaアームでは、他タスクの実デモで微調整したモデルが、単一の人間動画からコンパイルしたプログラムで40試行中36を完了(ビデオ最終フレームをゴール画像とした場合は22)。

3. 技術・手法の肝は?

- デモンストレーションをeffect programにコンパイル:動いた物体の3Dキーポイント軌跡、各物体の把持点を示す2点、デモンストレータを除いたシーンの終端配置。 - PEWAMは71.5Mパラメータの世界行動モデルで、effect、robot execution、action、terminal stateの4ストリームを独立したflow-matching時間で生成。 - プログラムをクランプし実行をサンプリングすることで、推論をプログラミングとして扱い、ライブシーンから閉ループで再解決。 - プログラムは座標の集合であるため、人間が編集可能(配置は終端状態のシフトに従い、把持は接触点の回転に応じて変化)。

4. どうやって有効だと検証した?

- LIBERO-Goalの未見タスクで、1デモのプログラムが90エピソード中40を完了。比較手法は最大19。 - Meta-Worldの未見3クラスで公開最高結果を上回る(5クラス平均では及ばない)。 - 同じプログラムを4つのロボットアームで再学習なしに実行。 - 押し動作後の再解決で60エピソード中23を完了(デモ再生は6)。 - Frankaアームで、他タスクの実デモで微調整したモデルが、単一の人間動画からコンパイルしたプログラムで40試行中36を完了(ビデオ最終フレームをゴール画像とした場合は22)。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- LIBERO-Goal、Meta-World、Frankaアームを用いた研究。 - デモンストレーション条件付き手法、ゴール画像条件付き手法、言語条件付き手法。 - flow-matching、world-action model、effect program。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junyi Hu, Zhewen He, Zhenhua Li, Yi Fang

分類: cs.RO

原文アブストラクト

A robot demonstration records two things in the same frames: what happened to the objects, and how one particular arm made it happen. We condition on the first. A demonstration is compiled into an effect program: the 3D keypoint trajectories of the objects that moved, two points marking where each was held, and the configuration the scene ends in, with the demonstrator removed. PEWAM, a 71.5M-parameter world-action model, generates effect, robot execution, action and terminal state as four streams with independent flow-matching times, so clamping a program and sampling the execution turns inference into programming, re-solved closed loop from the live scene. On held-out LIBERO-Goal tasks, one demonstration's program completes 40 of 90 episodes, where the same backbone given a goal image or language, and published demonstration-conditioned methods, complete at most 19; on three of Meta-World's held-out classes it exceeds the best published results, though not on the five-class mean. Because a program is a set of coordinates, a person can edit it: the placement follows a shifted terminal state and the grasp turns with rotated contact points. The same program runs on four robot arms without retraining, and after a push, re-solving completes 23 of 60 episodes where replaying the demonstration completes 6. On a Franka arm, fine-tuned on real demonstrations of other tasks, programs compiled from single human videos complete 36 of 40 trials, against 22 for the same backbone conditioned on the video's last frame as a goal image.

関連論文

PR本紙発行元 EmplifAI