日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.10855

OmniHOI: 単眼人間動画からの巧みな手と物体のインタラクション

OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video

シェア:XThreadsFacebookLINEはてブBluesky

単眼RGB動画から手と物体のインタラクションを再構成し、物理的一貫性を段階的に強制することで、多様な自由度のロボットハンドへ高成功率で転送するパイプラインを提案。

詳しい要約

1. どんなもの?

- 単眼RGB動画からdexterous handの手-物体interaction軌道を生成するパイプラインOmniHOIを提案。 - 再構成・retargeting・physics-in-the-loop refinementの各段階で物理整合性を強制。 - 150のmotion-capture軌道と60の単眼動画クリップで評価し、実機bimanual robotでも実行。

2. 先行研究と比べてどこがすごい?

- 従来はタスク特化RL訓練が必要でスケーラビリティに難、またはクリーンなmotion-capture軌道を仮定し動画に直接適用不可。 - OmniHOIは動画から直接、物理整合的な軌道を生成。 - 5種のdexterous hand(6-22 DoF)で39-89%成功、従来は最大31%。 - 単眼動画60クリップで53%成功、最良の従来video-to-robotパイプラインは28%。

3. 技術・手法の肝は?

- 各段階で利用可能な証拠に基づき物理整合性を強制。 - 再構成時はimage evidence、retargeting時はcontact geometry、refinement時はdynamicsを利用。 - 各段階で対応する表現を直接最適化し、誤差が下流に伝播する前に修正。

4. どうやって有効だと検証した?

- 150のmotion-capture軌道を5種のdexterous hand(6-22 DoF)に転送し、成功率39-89%を達成。 - 60の単眼動画クリップで53%成功。 - 実機bimanual robotで多様なタスクを実行。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている先行研究: task-specific RL training、motion-capture軌道を仮定する手法、video-to-robotパイプライン。 - 関連手法: dexterous hand retargeting、physics-in-the-loop refinement、hand-object interaction reconstruction。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ting Mao, Yanming Shao, Ziheng Wang, Haoyu Liu, Yiqun Wang, Xuanye Wu, Yao Mu

分類: cs.RO

原文アブストラクト

Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution. Prior methods either require task-specific RL training, limiting scalability, or assume clean motion-capture trajectories and thus cannot operate directly on video. We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands. The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement. Each stage optimizes the corresponding representation directly, correcting errors before they propagate downstream or must be absorbed by a learned policy. Across 150 motion-capture trajectories transferred to each of five dexterous hands with 6 to 22 DoF, we achieve 39-89% success, compared with at most 31% for prior transfer methods. On 60 monocular video clips, we achieve 53% success, compared with 28% for the best prior video-to-robot pipeline. Its trajectories also execute on a real bimanual robot across diverse tasks.

関連論文

PR本紙発行元 EmplifAI