分解と再構成:既存能力から新スキルを推論するクロスタスクロボット操作
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation
見たタスクのデモを原子スキルと行動のペアに分解し、未見タスクに対してスキルを再構成して推論するフレームワークを提案。ゼロショットのクロスタスク汎化を実現した。
著者: Xitie Zhang, Aming Wu, Yahong Han
分類: cs.RO, cs.CV
原文アブストラクト
Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge from seen tasks. Recent in-context learning approaches leverage seen task demonstrations to generate actions for unseen tasks without parameter updates. However, existing methods provide only low-level continuous action sequences as context, failing to capture composable skill knowledge and causing models to degenerate into superficial trajectory imitation. We propose Decompose and Recompose, a skill reasoning framework using atomic skill-action pairs as intermediate representations. Our approach decomposes seen demonstrations into interpretable skill--action alignments, enabling the model to recompose these skills for unseen tasks through compositional reasoning. Specifically, we construct a task-adaptive dynamic demonstration library via visual-semantic retrieval combined with skill sequences from a planning agent, complemented by a coverage-aware static library to fill missing skill patterns. Together, these yield skill-comprehensive demonstrations that explicitly elicit compositional reasoning for skill composition and execution ordering. Experiments on the AGNOSTOS benchmark and real-world environments validate our method's zero-shot cross-task generalization capability.
関連論文
- 意図を考慮したロボットから人への両手受け渡しのための時間的触覚符号化とコンプライアンス制御マニピュレーション
- ロボットハンド操作における形態と駆動方式の帰納的バイアスマニピュレーション
- 動作中の着衣支援:人間の動きを考慮した拡散ポリシーによるロボット着衣支援マニピュレーション
- 適応的視覚言語把持:構成可能な基盤事前知識と汎化可能な把持合成による実現マニピュレーション
- GIFT: 行動指向の構造的監督による誘導中間特徴学習を用いたロボット操作マニピュレーション
- HINT: 長期的ロボット操作のための人間意図の注入マニピュレーション