日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
全身マニピュレーションarXiv:2609.30735

Praxis: 一人称視点動画から物理的相互作用の事前知識を蒸留し、汎用的な全身マニピュレーションを実現

Praxis: Distilling Physical Interaction Priors from Egocentric Videos for Generalizable Whole-Body Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

一人称視点の実演動画から物理的相互作用の事前知識を抽出し、視覚言語誘導ナビゲーション・姿勢キャリブレーション・器用な全身操作を組み合わせることで、タスクごとの再学習なしに多様な物体操作を可能にするフレームワークを提案。

詳しい要約

1. どんなもの?

- モバイルヒューマノイドの全身マニピュレーションを実現するフレームワーク「Praxis」を提案。 - 一人称視点のワンショット動画デモから物理的相互作用の事前分布を抽出し、閉ループ姿勢調整とオンライン知覚を組み合わせる。 - 視覚言語ガイドによるナビゲーション、腕-手作業空間を合わせる閉ループ姿勢調整、上下半身を同期させた器用な操作の3段階を統合。 - オンライン視覚フィードバックで新しい物体姿勢やシーン構成に再接地し、触覚フィードバックで接触条件に適応。 - 各操作スキルは人間のデモ1回で指定され、タスク固有の操作ポリシー再学習を不要とする。

2. 先行研究と比べてどこがすごい?

- 従来は限られたタスク固有データから全身操作を学習するのが困難だった。 - Praxisはワンショットの一人称動画デモから物理的相互作用の事前分布を蒸留し、タスク固有ポリシーの再学習なしで多様な操作を実現。 - 空間的・視覚的・物体横断的な汎化と、3段階すべてでの外部物理的擾乱からの回復を可能にする点が優位。 - 具体的な比較対象は要旨からは不明。

3. 技術・手法の肝は?

- 一人称動画デモから物理的相互作用の事前分布を抽出。 - 3段階の調整:視覚言語ガイドによるナビゲーション、閉ループ姿勢調整による腕-手作業空間の整合、上下半身同期制御による器用な操作。 - オンライン視覚フィードバックでデモの相互作用幾何を新しい物体姿勢・シーン構成に再接地。 - 触覚フィードバックで実際の接触条件に手の動きを適応。 - 各操作スキルは人間のデモ1回で指定され、タスク固有の操作ポリシー再学習を必要としない。

4. どうやって有効だと検証した?

- 5つの長期的操作タスクで実験を実施。 - 空間的汎化、視覚的汎化、物体横断的汎化を実証。 - 3段階すべてにおいて外部物理的擾乱からの回復を確認。 - 具体的な評価指標やベースラインは要旨からは不明。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、実世界展開の課題などについての議論は記述されていない。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない。 - 同分野の関連手法として、一人称動画からの模倣学習(例:egocentric video imitation learning)、視覚言語モデルを用いたナビゲーション(例:vision-language navigation)、触覚フィードバックに基づくマニピュレーション(例:tactile manipulation)、全身制御(例:whole-body control)などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shuliang He, Ruiyan Xu, Bo Yue, Hengming Zhang, Huayi Zhou, Shuai Wang, Wei-Shi Zheng, Guiliang Liu

分類: cs.RO

原文アブストラクト

Mobile humanoid manipulation requires both reaching a usable workspace and preserving precise hand-object interactions as object poses and contact conditions change. Learning these behaviors from limited task-specific data remains challenging. To bridge this gap, we introduce Praxis, a whole-body manipulation framework that combines physical interaction priors from one-shot egocentric video demonstrations with closed-loop posture calibration and online perception. The framework coordinates three stages: vision-language-guided navigation toward target objects, closed-loop posture calibration to align the arm-hand workspace, and dexterous manipulation with synchronized upper- and lower-body control. Online visual feedback re-grounds demonstrated interaction geometry under new object poses and scene configurations, while tactile feedback adapts hand motions to actual contact conditions. Each manipulation skill is specified by one human demonstration, without task-specific manipulation-policy retraining. Experiments across five long-horizon manipulation tasks demonstrate spatial, visual, and cross-object generalization, as well as recovery from external physical disturbances across all three stages.

関連論文

PR本紙発行元 EmplifAI