日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.17714

視覚言語に基づくタスク文脈認識型模倣学習によるロボット分解

Vision-Language Grounded Task-Context-Aware Imitation Learning for Robotic Disassembly

シェア:XThreadsFacebookLINEはてブBluesky

言語指示を視覚空間表現に接地させ、階層的なタスク選択とタスク文脈認識型模倣学習を組み合わせることで、多様な形状や配置の部品を対象としたロボット分解の長期的タスク遂行を実現した。

詳しい要約

1. どんなもの?

- ロボットによる分解作業(robotic disassembly)を対象とした模倣学習フレームワーク。 - 長期的な作業(long-horizon execution)で、複数部品に対する順序付き操作を必要とする。 - 言語によるタスクコンテキストを導入し、階層的タスク選択とタスクコンテキスト認識型模倣学習を組み合わせる。 - 言語指示を視覚空間表現にグラウンディングし、多様なコネクタ形状や組立構成に汎化する。 - 明示的な物体アノテーションを必要としない。

2. 先行研究と比べてどこがすごい?

- 従来の模倣学習(例:diffusion policy)やタスクコンテキスト認識ベースラインと比較。 - エンドツーエンドのタスク成功率がdiffusion policyベースラインより35ポイント、以前のタスクコンテキスト認識ベースラインより75ポイント向上。 - 生の観測のみから意図されたスキルを推論する困難さを、言語によるタスクコンテキストで緩和。 - 訓練データがカバーできない組み合わせ的多様性に対処できる点が優位。

3. 技術・手法の肝は?

- 階層的タスク選択(hierarchical task selection)とタスクコンテキスト認識型模倣学習(task-context-aware imitation learning)を統合。 - 言語指示を空間視覚表現にグラウンディングし、言語指定タスクと対応する操作対象を視覚シーン内で関連付ける。 - タスクコンテキストがスキル選択の明示的な構造を提供。 - 明示的な物体アノテーションを不要とする。

4. どうやって有効だと検証した?

- エンドツーエンドのタスク成功率を指標に評価。 - ベースライン(diffusion policy)および以前のタスクコンテキスト認識ベースラインと比較。 - 多様なコネクタ形状や組立構成への汎化を検証。 - 具体的な実験設定やデータセットは要旨からは不明。

5. 議論はある?

- 言語によるタスクコンテキストがスキル選択と操作対象の関連付けに有効であることを示唆。 - 訓練データの組み合わせ的多様性の限界を言語が補う可能性。 - 限界や失敗事例、計算コスト、実世界への適用範囲については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されているベースライン:diffusion policy、以前のタスクコンテキスト認識ベースライン。 - 関連手法:imitation learning、vision-language grounding、hierarchical task selection。 - 同分野の定番:behavior cloning、transformer-based policies。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jeon Ho Kang, Igal Tamarkin, Ethan Niu, Ian Novales, Satyandra K. Gupta

分類: cs.RO

原文アブストラクト

Real-world robotic disassembly requires long-horizon execution, where robots must perform ordered sequences of manipulation tasks across multiple parts within a single scene. Multiple valid task goals and diverse assembly configurations make it difficult for imitation policies to infer the intended skill from raw observations alone, particularly when training data cannot cover the combinatorial diversity of real-world configurations and part geometries. We show that incorporating task context through language alleviates these challenges by providing explicit structure for skill selection and associating language-specified tasks with their corresponding manipulation targets in the visual scene. The proposed framework combines hierarchical task selection with task-context-aware imitation learning to ground language instructions in spatial visual representations for robotic disassembly. The resulting framework generalizes across diverse connector geometries and assembly configurations without requiring explicit object annotations. Our method improves end-to-end task success by 35 percentage points over the baseline diffusion policy and by 75 percentage points over the previous task-context-aware baseline.

関連論文

PR本紙発行元 EmplifAI