日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.19138

VLMエージェントによる文脈内ロボット学習

In-Context Robot Learning with VLM Agents

シェア:XThreadsFacebookLINEはてブBluesky

商用VLMを活用し、デモやフィードバックから勾配更新なしでロボット行動を生成するGPT-Policyを提案。実機実験で人間のビデオデモがタスク完了率を向上させることを示した。

詳しい要約

1. どんなもの?

- 商用VLM(例:GPT-6 Astra)のエージェント能力を活用し、勾配更新やタスク固有パラメータの永続的変更なしに、実ロボットが文脈から学習する枠組み「GPT-Policy」を提案。 - 文脈コンパイラ、ロボットツール行動を提案するVLM、行動を検証・実行し結果を報告する制約付きコントローラで構成。 - 実ロボット試行で、人間のビデオデモンストレーションがロボット行動ラベルなしでもタスク完了を改善し、接触に敏感なタスクでは整列した行動参照がさらなる向上をもたらすことを示す。

2. 先行研究と比べてどこがすごい?

- 既存のロボティクスポリシーはin-context learning(ICL)がほぼ不可能であり、有限のデモンストレーションでは未知のタスクや状況に対応できない。 - 本研究は商用VLMの広範なエージェント能力を活用し、展開時の文脈から学習する能力をロボットにもたらす点で先行研究と異なる。 - 勾配更新やタスク固有パラメータの永続的変更を必要とせず、新たな初期状態から実行可能で検証可能なロボット行動を生成できる可能性を示した。

3. 技術・手法の肝は?

- 文脈コンパイラ:タスクに関連する視覚的遷移を保持・抽出。 - VLM:ロボットツール行動を提案。 - 制約付きコントローラ:各行動を検証・実行し、その結果を報告。 - これらを統合し、勾配更新なしで文脈から学習する枠組みを実現。

4. どうやって有効だと検証した?

- タスク成功率と効率指標、モデル間のマッチ比較、制御された文脈アブレーションを通じて信頼性と限界を評価。 - 実ロボット試行を実施し、人間のビデオデモンストレーションがロボット行動ラベルなしでタスク完了を改善することを確認。 - 接触に敏感なタスクでは、整列した行動参照がさらなる性能向上をもたらすことを示した。

5. 議論はある?

- 本研究は、VLMの汎用能力を物理行動に変換するための実証的基盤を提供し、信頼性の高い展開に向けた課題を明らかにする。 - ただし、具体的な限界や課題の詳細は要旨からは不明。 - 実ロボット試行の結果は示されているが、議論の深掘りは要旨では触れられていない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、in-context learning (ICL)、vision-language models (VLMs)、GPT-6 Astra、ロボティクスポリシー、制約付きコントローラなどが挙げられる。 - 同分野の定番として、模倣学習、強化学習、大規模言語モデルを活用したロボット学習に関する論文が次に読むべき候補となる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu

分類: cs.CV, cs.RO

原文アブストラクト

Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.

関連論文

PR本紙発行元 EmplifAI