日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.06843

エージェント型ロボットのための再帰的ビデオ文脈内学習

Recursive Video In-Context Learning for Agentic Robot

シェア:XThreadsFacebookLINEはてブBluesky

実演ビデオをサブイベント階層として構築し、エージェントが必要な時に必要な粒度だけを読み込む訓練不要の手法を提案。LIBERO-PROとLIBERO-Plusで成功率を向上させた。

詳しい要約

1. どんなもの?

LLM agent が frozen な VLA policy を編成する際、text memory では「何をしたか」は記録できるが「どうやるか」が欠落する問題に対し、demonstration video を階層構造として扱う training-free 手法 RV-ICL を提案。 - demonstration を prompt として与えるのではなく、agent が navigation する hierarchy に変換。 - hierarchy は grasp や release などの sub-event から構築。 - 全体 keyframe から phase、moment、short clip へと階層が細かくなる。 - read-only tool 経由で coarse level を planning 前に読み、実行中は必要時に再進入し現 sub-goal の clip のみ load。 - 1 task あたり demonstration 1 本で十分。

2. 先行研究と比べてどこがすごい?

従来の LLM agent は text memory で episode を跨ぎ改善するが、記録は行動履歴に留まり task の遂行方法(how)を持たない。 - demonstration video は how を示すが、agent の context に収まりにくい。 - full video は毎 turn を遅くし、fixed keyframe は grasp 成否を決める contact 詳細を失う。 - 必要な情報は planning 時の task 構造から、各 contact 周辺 frame へと shift する。 - RV-ICL は video を prompt ではなく階層 navigation 対象とし、必要部分のみ読む点が先行手法と異なる。

3. 技術・手法の肝は?

training-free で demonstration を階層化し、agent が read-only tool で辿る仕組み。 - demonstration の sub-event(grasp, release 等)から hierarchy を構築。 - 階層は whole task の keyframe → phase → moment → short clip と細かくなる。 - planning 前に coarse level を読み、実行中に step が詳細を要する時のみ hierarchy へ再進入。 - 現 sub-goal の clip だけを load し context を抑制。 - 1 task 1 demonstration で動作。

4. どうやって有効だと検証した?

RPent を基盤に LIBERO-PRO と LIBERO-Plus で評価。 - LIBERO-PRO で success が 92.6% から 96.5% へ向上。 - LIBERO-Plus で 86.7% から 95.8% へ向上。 - 1 task あたり demonstration 1 本で達成。 - 具体的な ablation や比較条件は要旨からは不明。

5. 議論はある?

要旨からは不明。 - 想定される論点として full video の遅さ、fixed keyframe の contact 詳細欠落、必要情報の planning から contact 周辺への shift が挙げられている。 - 階層構築のコストや失敗時の挙動、他 task への汎化は要旨では触れられていない。

6. 次に読むべき論文は?

要旨で参照/比較されている研究・手法。 - RPent(基盤 agent) - frozen VLA policies - LIBERO-PRO - LIBERO-Plus - text memory を用いる LLM agents - demonstration video を prompt として扱う手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang

分類: cs.RO, cs.AI, cs.CL, cs.MA

原文アブストラクト

LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.

関連論文

PR本紙発行元 EmplifAI