日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.35375

ピクセルからポーズへ:人間のデモンストレーションによる物体中心のツール操作学習

From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

シェア:XThreadsFacebookLINEはてブBluesky

人間の動画デモから物体中心の世界モデルを事前学習し、ポーズ認識方策と組み合わせることで、ロボットとの対応データなしに複雑なツール操作を実現した。

詳しい要約

1. どんなもの?

- 人間のデモンストレーションから直接ツール操作を学習するデータ効率の良い物体中心フレームワーク「P2P-T」を提案。 - 2段階アプローチ:物体中心のworld modelを事前学習して安定したpose priorを抽出し、それを効率的なpose-aware低レベルポリシーに統合。 - 現代のfoundation modelを活用した自動データ処理パイプラインにより、人間とロボットのアラインメントデータを完全に不要に。 - 複雑な実世界ツール操作タスクにおいて、最小限のタスクごとのfine-tuningで従来のstate of the artを73%上回る実行性能を達成。

2. 先行研究と比べてどこがすごい?

- 従来の人間動画デモ利用手法は計算コストが高く、ドメインアラインメントのために人間-ロボットのペアデータに依存。 - 既存のstate of the artは長期的タスクは得意だが、複雑なツール操作に必要な繊細で精密な制御は苦手。 - P2P-Tは人間-ロボットアラインメントデータを完全に不要にし、訓練オーバーヘッドを大幅削減。 - 複雑な実世界ツール操作タスクで従来のstate of the artを73%改善し、標準的な大規模事前学習モデルでは到達できない性能を実現。

3. 技術・手法の肝は?

- 2段階アプローチ:第一に物体中心のworld modelを事前学習し、安定したpose priorを抽出。 - 第二にこれらのpriorを効率的なpose-aware低レベルポリシーに統合。 - 現代のfoundation modelを活用した堅牢な自動データ処理パイプラインを利用。 - これにより人間-ロボットアラインメントデータを完全にバイパスし、訓練オーバーヘッドを大幅に削減。

4. どうやって有効だと検証した?

- 複雑な実世界ツール操作タスクにおいて、最小限のタスクごとのfine-tuningで実行性能を評価。 - 従来のstate of the artと比較して73%の改善を達成。 - 標準的な大規模事前学習モデルでは到達できないタスクで有効性を検証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:人間動画デモを利用する最近のアプローチ、長期的タスクに優れる現在のstate of the art、標準的な大規模事前学習モデル。 - 関連手法:object-centric world model、pose-aware low-level policy、foundation modelを活用したデータ処理パイプライン。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bangjun Wang, Longyan Wu, Yukun Wei, Shenghe Shao, Chaoyi Huang, Wenze Cui, Zetong Xu, Hanlin Wu, Long Chen, Yi Ma, Hongyang Li

分類: cs.RO, cs.AI

原文アブストラクト

Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.

関連論文

PR本紙発行元 EmplifAI