日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
リターゲティングarXiv:2609.37776

単眼動画からの上半身人体-ロボット動作リターゲティングにおける幾何構造保持手法

Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video

シェア:XThreadsFacebookLINEはてブBluesky

単眼RGB動画から人体の上半身と手の動きを統一的に復元し、形態に依存しない幾何学的変換と多段階逆運動学を用いてロボットへ転移するフレームワークを提案。

詳しい要約

1. どんなもの?

- 単眼RGB動画から上半身の人間動作をロボットへ転移する枠組み - 身体と手の統合再構成と形態非依存の幾何転移を組み合わせる - 対象はdual-arm dexterous robotの上半身協調動作 - Across-VAMが動画生成、Across-WAMがhuman-to-robot motion mappingを担う - 単眼動画からロボット上半身動作への一貫パイプラインを提示

2. 先行研究と比べてどこがすごい?

- 従来は身体と手を別スケールで復元し転移が困難だった - 本手法は統合再構成でhand reprojection errorを22.36から7.21 pixelsへ低減 - SAM 3D Body比で手の再投影誤差が改善 - 形態差が大きいhumanとrobot間でもfine distal motionを保持 - 単眼動画からdual-arm dexterous robotへの協調動作を実現

3. 技術・手法の肝は?

- frame-wise body estimatesとvideo-level observations、hand evidenceを統合 - 単一のdifferentiable Momentum Human Rig (MHR) stateで拘束 - 一過性のhand artifactsをparameter spaceで修復 - 動作をarm-segment directions、elbow configuration、relative palm orientation、bilateral wrist relationsで表現 - multi-stage inverse kinematicsとrobot-specific hand adaptationで実現

4. どうやって有効だと検証した?

- 16本の単眼動画、計1,769 source framesで評価 - 内訳はsigning 10本、reach-to-grasp 6本 - 統合再構成によりhand reprojection errorが22.36から7.21 pixelsへ低減 - 16本すべてのretargeted trajectoriesがkinematic simulation playbackを完了 - 代表的なsigningとreach-to-grasp動作をphysical robotで実演

5. 議論はある?

- 単眼動画からdual-arm dexterous robotへの統合パイプラインを実証 - 身体と手の空間スケール差、運動学的差異、fine distal motion保持が課題 - 一過性のhand artifactsをparameter spaceで修復する点に言及 - 評価は16動画・1,769フレームと限定的 - 実機実演は代表動作のみで、一般化や失敗例の議論は要旨からは不明

6. 次に読むべき論文は?

- SAM 3D Body(比較対象として参照) - Momentum Human Rig (MHR)(基盤表現として参照) - Across-VAM(動画生成モジュールとして参照) - Across-WAM(human-to-robot motion mappingとして参照) - 同分野の定番としてhuman-to-robot motion retargetingおよびinverse kinematics関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiaoyu Yang, Sen Han, Da Li, Nan Wu

分類: cs.RO

原文アブストラクト

Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer. Frame-wise body estimates, video-level observations, and detailed hand evidence jointly constrain a single differentiable Momentum Human Rig (MHR) state, while transient hand artifacts are repaired in parameter space. The reconstructed motion is represented by arm-segment directions, elbow configuration, relative palm orientation, and bilateral wrist relations, and is realized on the target robot through multi-stage inverse kinematics and robot-specific hand adaptation. Within the broader system, Across-VAM provides video generation, whereas Across-WAM performs human-to-robot motion mapping. The method is evaluated on 16 monocular videos comprising 1,769 source frames, including 10 signing and six reach-to-grasp sequences. Unified reconstruction reduces mean hand reprojection error from 22.36 to 7.21 pixels relative to SAM 3D Body. All 16 retargeted trajectories completed kinematic simulation playback, and representative signing and reach-to-grasp motions were further demonstrated on a physical robot. The results demonstrate a unified pipeline from monocular human video to coordinated upper-body motion on a dual-arm dexterous robot.

関連論文

PR本紙発行元 EmplifAI