日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.37359

ENCORE: 少数の実演から操作戦略を発見するエージェント

Encore: Few-Shot Agentic Discovery of Manipulation Strategies

シェア:XThreadsFacebookLINEはてブBluesky

少数の実演を証拠として与え、コーディングエージェントが知覚・行動APIに対するポリシープログラムを書いて反復改善し、言語指示だけでは不明な操作戦略を発見する手法。

詳しい要約

1. どんなもの?

ENCOURAGEは、coding agentが少数のdemonstrationを「訓練データ」ではなく「読むべき証拠」として与えられ、ロボット操作戦略を発見するfew-shot agentic discovery手法。 - 対象: 文だけではgraspや接触順序、結果が不明なrobot tasks。 - 構成: deterministic builderが各demonstrationをmulti-view keyframes、gripper events、frame strips、full trajectoryのpackに蒸留。 - 流れ: coding agentがpackを読み、固定のperception/action APIに対してpolicy programを書き、数回のdevelopment rolloutsで反復改善し、成功信号を見ないsealed evaluation前にfreeze。 - 評価: LIBERO-PRO、RoboDojo、実機bimanual robot。

2. 先行研究と比べてどこがすごい?

先行のagentic systemと比べ、demonstrationを訓練データでなく証拠として読ませる点が新しい。 - 同じlanguage modelを使った最強のprior agentic systemに対し、frozen programsが96.3% vs 89.3%で上回る。 - LIBERO-PROのperturbed tasksで、agentの最初のprogramがdemonstrationありで半数成功、demonstrationなしでは1タスクのみ成功。 - RoboDojoのgoalが明示されないinstructionsでは、demonstrationなしではどのprogramも成功しない。

3. 技術・手法の肝は?

肝はdemonstrationを構造化packに変換し、coding agentがAPI準拠のpolicy programを書いて反復改善する点。 - deterministic builder: multi-view keyframes、gripper events、frame strips、full trajectoryを生成。 - coding agent: 固定のperception and action APIに対してpolicy programを作成。 - 改善: 少数のdevelopment rolloutsでiterative refinement。 - 評価: 成功信号を明かさないsealed evaluationの前にprogramをfreeze。

4. どうやって有効だと検証した?

LIBERO-PRO、RoboDojo、実機bimanual robotで検証。 - LIBERO-PRO: 最初のprogramがdemonstrationありでperturbed tasksの半数成功、なしでは1タスク成功。 - 比較: 同じlanguage modelの最強prior agentic systemに96.3% vs 89.3%で勝利。 - RoboDojo: goal未記載のinstructionsではdemonstrationなしで成功programなし。 - 実機: 5 demonstrationsずつからcube handoverとcup inversionを学習。

5. 議論はある?

要旨からは、限界や失敗事例、計算コスト、demonstration品質への依存性についての議論は不明。 - 主張: demonstrationを証拠として読むことでfew-shot discoveryが可能。 - 未解決: sealed evaluationの詳細、成功信号の定義、一般化範囲は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究や関連手法を挙げる。 - 最強のprior agentic system(同一language model使用、名称は要旨からは不明)。 - LIBERO-PRO、RoboDojo。 - coding agents、robot manipulation、few-shot learning、bimanual robot manipulationの同分野定番。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yifan Kang, Zihan Wang, Zhiwen Fan, Bangya Liu

分類: cs.RO, cs.AI, cs.CV, cs.LG

原文アブストラクト

Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We introduce ENCORE, which gives the agent a few demonstrations as evidence to read rather than as training data. A deterministic builder distills each demonstration into a pack of multi-view keyframes, gripper events, frame strips, and the full trajectory. A coding agent studies the pack, writes a policy program against a fixed perception and action API, refines it iteratively over a few development rollouts, and freezes it before a sealed evaluation that never reveals the success signal. On LIBERO-PRO, the agent's first program already succeeds in half of the perturbed tasks with demonstrations and in one task without them, and the frozen programs outperform the strongest prior agentic system run with the same language model (96.3% against 89.3%). On RoboDojo tasks whose instructions leave the goal unstated, no program succeeds without demonstrations. ENCORE also runs on a real bimanual robot, learning cube handover and cup inversion from five demonstrations each.

関連論文

PR本紙発行元 EmplifAI