日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.27780

タスクプロトタイプ誘導フローマッチングによる視覚言語ロボット操作の少数ショット汎化

Task-Prototype Guided Flow Matching for Few-Shot Generalization in Vision-Language Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

少数の実演からタスクの位相レベルのプロトタイプを抽出し、フローマッチングの初期分布と速度場を誘導することで、視覚言語ロボット操作を新しい手順に少数ショットで適応させる手法を提案。

詳しい要約

1. どんなもの?

- 視覚言語ロボット操作ポリシーを少数のデモから新しい手順に適応させるフレームワーク。 - 言語では不十分な接触タイミング、動作フェーズ、修正行動、実行スタイルを補う。 - サポートデモを構造化されたタスクプロトタイプトークンに変換し、フローマッチングをガイド。 - 1-shot、4-shot、6-shot設定で成功率66.8%、79.6%、82.1%を達成。

2. 先行研究と比べてどこがすごい?

- 1-shotでCFM、Pooled-Demo CFM、In-Context Flowをそれぞれ29.8、14.5、9.9ポイント上回る。 - 新規物体転移、目標再結合、長地平線構成、接触/修正タスクでの汎化も改善。 - ノイズの多いサポートでの成功率低下を6.2%に抑える。

3. 技術・手法の肝は?

- 対称クロスアテンションと学習可能クエリでフェーズレベルのプロトタイプを抽出。 - タスク適応型初期分布をパラメータ化し、ゲート付き適応正規化でプロトタイプ情報を注入。 - エピソード的サポート・クエリ目的とプロトタイプ対照正則化で訓練。 - 訓練中に少数ショット適応をシミュレートし、不要情報を抑制。

4. どうやって有効だと検証した?

- LEROBOT-ARM-SO101プラットフォームで1-, 4-, 6-shot設定の成功率と75.5%の少数ショットAUCを評価。 - 新規物体転移、目標再結合、長地平線構成、接触/修正タスクでの汎化を検証。 - リアルタイム実行性能(6オンラインプロトタイプトークン、54.3ms遅延、3.9GBピークメモリ、10Hz制御)を測定。 - 理論的診断でプロトタイプ距離と行動分布距離の整合性、適応事前分布の輸送コスト削減、ゲート付き変調のODE境界内抑制を確認。

5. 議論はある?

- 理論的診断はプロトタイプ距離と行動分布距離の整合性、適応事前分布の輸送コスト削減、ゲート付き変調のODE境界内抑制を示す。 - ノイズの多いサポートでの成功率低下が6.2%に抑えられることを確認。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- CFM - Pooled-Demo CFM - In-Context Flow - 同分野の定番としてFlow Matching、Vision-Language Robot Manipulation、Few-Shot Imitation Learning

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yizhao Wang, Guantao Zhang, Jingbo Wang

分類: cs.RO, cs.CV

原文アブストラクト

Vision-language robot manipulation policies can follow semantic instructions, but adapting them to a new procedure from only a few demonstrations remains difficult because language underspecifies contact timing, motion phases, corrective behavior, and execution style. This paper presents Task-Prototype Guided Flow Matching (TP-Flow), a few-shot manipulation framework that converts support demonstrations into structured task-prototype tokens and uses them to guide both the initial flow prior and the velocity field. TP-Flow employs symmetric cross-attention with learnable queries to extract phase-level prototypes, parameterizes a task-adaptive initial distribution, and injects prototype information through gated adaptive normalization. It is trained with an episodic support-query objective and prototype contrastive regularization, so few-shot adaptation is simulated during training while nuisance information is suppressed. On the LEROBOT-ARM-SO101 platform, TP-Flow achieves 66.8\%, 79.6\%, and 82.1\% success rates under 1-, 4-, and 6-shot settings, with a 75.5\% few-shot AUC. At 1-shot, it improves over CFM, Pooled-Demo CFM, and In-Context Flow by 29.8, 14.5, and 9.9 percentage points. It also improves held-out target-group generalization across novel-object transfer, goal recombination, long-horizon composition, and contact/correction tasks. TP-Flow maintains real-time execution with six online prototype tokens, 54.3 ms latency, 3.9 GB peak memory, and a 10 Hz control rate, while reducing the noisy-support success drop to 6.2\%. Theoretical diagnostics show that prototype distance aligns with action-distribution distance, the adaptive prior reduces transport cost, and gated modulation keeps measured trajectory deviations below the derived ODE bound. The code repository is omitted for anonymous review.

関連論文

PR本紙発行元 EmplifAI