日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.28314

TANDEM: 必要時デモンストレーションを用いたタスク・動作計画による視覚言語行動モデルの効率的ファインチューニング

TANDEM: Task and Motion Planning with As-Needed Demonstrations for Efficient Vision-Language-Action Model Fine-tuning

シェア:XThreadsFacebookLINEはてブBluesky

タスク・動作計画(TAMP)と必要な時だけ人間が遠隔操作する手法を組み合わせ、長期的なマニピュレーションタスクのデモンストレーションを効率的に収集し、視覚言語行動モデルのファインチューニングを可能にするシステムを提案。

詳しい要約

1. どんなもの?

- TANDEMは、Task and Motion Planning (TAMP) と選択的な人間のteleoperationを組み合わせたシステム。 - 言語指示と視覚観測から、pretrained vision-language modelsを用いて計画ドメインを拡張し、不足するpredicatesと人間が実行するmagic operatorsを追加。 - これにより、タスク固有の介入点なしに自律段階と人間実行段階を交互に実行可能。 - 人間段階の後、シーンを再知覚し意図した効果を確認してから自律計画を再開。 - VLAモデルのfine-tuningのため、example pretraining trajectoriesを用いてplanner生成動作を対象モデルのpretraining分布に整合させる。

2. 先行研究と比べてどこがすごい?

- 従来のteleoperationは、ロボットが自律的に実行できる行為にも多くの人間デモを要し、データ収集のスケーラビリティを制限。 - 固定されたTAMPドメインでは長期的操作タスクの全段階を支援できない。 - TANDEMは、人間の支援をオンデマンドの計画能力として表現し、TAMPドメインを動的に拡張。 - タスク固有の介入点を必要とせず、自律と人間実行を交互に組み合わせられる点が新しい。 - 同じ人間介入時間で、full-task teleoperationの2.9倍のデモを収集可能。

3. 技術・手法の肝は?

- 言語指示と視覚観測を入力とし、pretrained vision-language modelsを用いてTAMPドメインに不足するpredicatesとhuman-executed magic operatorsを追加。 - これによりplannerが自律段階と人間実行段階をタスク固有の介入点なしに交互に計画。 - 各人間段階の後、シーンを再知覚し、意図した効果が成立するか確認してから自律計画を再開。 - VLAモデルfine-tuningのため、example pretraining trajectoriesでplanner生成動作を対象モデルのpretraining分布に整合。 - 具体的なVLM名や整合手法の詳細は要旨からは不明。

4. どうやって有効だと検証した?

- TAMPドメインの能力を超える5つの長期的操作タスクで評価。 - 代表的な長期的タスクにおいて、同じ人間介入時間でfull-task teleoperationの2.9倍のデモを収集。 - pretrained π_{0.5}-DROIDモデルを各タスク20個のTANDEMデモでfine-tuning。 - 5タスク平均のタスク成功率が0%から60%に向上。 - これにより有効性を検証。

5. 議論はある?

- 要旨からは、TANDEMの限界や失敗事例、議論の詳細は不明。 - 人間介入時間やデモ数、成功率の向上は示されているが、他のベースラインとの比較や統計的有意性は要旨からは不明。 - 一般化可能性やスケーラビリティに関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: full-task teleoperation、Task and Motion Planning (TAMP)、vision-language-action (VLA) models、π_{0.5}-DROID。 - 関連手法として、pretrained vision-language modelsを用いた計画ドメイン拡張や、example pretraining trajectoriesによる動作整合。 - 同分野の定番として、Robot Learning from Demonstrations、Vision-Language-Action Models、Task and Motion Planningの基礎文献を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Samrat Sahoo, Liang Ji, Tom Silver, Yixuan Huang

分類: cs.RO

原文アブストラクト

Human teleoperators spend substantial time demonstrating behaviors that robots can already perform autonomously, limiting the scalability of data collection for robot foundation models. Task and motion planning (TAMP) can automate many of these behaviors, but a fixed planning domain may not support every stage of a long-horizon manipulation task. We present TANDEM (Tamp with As-Needed Demonstrations for Efficient Model fine-tuning), a system that combines TAMP with selective human teleoperation to collect demonstrations for tasks beyond the planner's capabilities. Our key idea is to represent human assistance as an on-demand planning capability. Given a language instruction and visual observation, TANDEM uses pretrained vision-language models to extend the planning domain with missing predicates and human-executed magic operators. This allows the planner to interleave autonomous and human-executed stages without task-specific intervention points. After each human stage, TANDEM re-perceives the scene and checks whether the intended effects hold before resuming autonomous planning. To support fine-tuning vision-language-action (VLA) models, TANDEM also uses example pretraining trajectories to align planner-generated motions with the target model's pretraining distribution. We evaluate TANDEM on five long-horizon manipulation tasks beyond the TAMP domain's capabilities. On a representative long-horizon task, TANDEM collects 2.9x as many demonstrations as full-task teleoperation at the same human intervention time. Fine-tuning a pretrained π_{0.5}-DROID model on 20 TANDEM demonstrations per task increases average task success from 0% to 60% across the five tasks.

関連論文

PR本紙発行元 EmplifAI