日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.24093

接触アンカリング再ターゲティングと残差方策学習による人間実演からの器用なロボットマニピュレーション

Dexterous Robot Manipulation from Human Demonstrations via Contact-Anchored Retargeting and Residual Policy Learning

シェア:XThreadsFacebookLINEはてブBluesky

人間の動作キャプチャから接触構造を抽出し、物理的に整合する軌道へ変換してロボットハンドの器用な操作方策を学習する3段階パイプラインを提案した。

詳しい要約

1. どんなもの?

- 人間のmotion-capture記録からdexterous robot policiesを学習する手法。 - 実ロボット訓練データを一切使わない。 - 3段階pipeline: physics refinement、contact-anchored retargeting、residual policy learning。 - 単一のresidual RL policyを多様なdemonstrationsで一度訓練し、物理的に整合したcontact-annotated trajectoriesへ修復。 - 25,454 single-hand trajectoriesと25 dual-hand tasksを再構成。 - 4つの形態の異なるrobot handsへ1つのhuman datasetを転移。 - 実機で4つのcontact-rich bimanual tasksを実行。

2. 先行研究と比べてどこがすごい?

- 従来のdemonstration学習は、grasp成功を決めるcontact forcesがscalableなhuman demonstrationsに欠けることがbottleneck。 - 本手法は、human handからrobot handへ残るのはjoint motionではなくcontact structure(どの指領域がどの物体位置に触れ、どの順序か)であると観察。 - 物理的整合性をタスクごとに設計する必要はなく、単一のresidual RL policyで多様なdemonstrationsを修復可能と主張。 - 同じresidual formulationがretargeting後の動的実現可能性も回復。 - 実ロボット訓練データなしでdexterous policiesを生成。 - 成功率: single-hand 7.3%→59.3%、dual-hand 16.0%→62.4%。 - 1つのhuman datasetを4つの形態の異なるrobot handsへ転移し+62.4 pp。

3. 技術・手法の肝は?

- 3段階pipeline。 - 第1段階: simulated MANO handによるphysics refinementでcontactsとforcesを回復。 - 第2段階: contact-anchored retargeting。hand morphologyに依存しないobjectiveで、demonstrated contact structureを転移。 - 第3段階: residual policy learningでrobot actuationに適応。 - 単一のresidual RL policyを多様なdemonstrationsで一度訓練し、kinematic recordingsを物理的に整合したcontact-annotated trajectoriesへ修復。 - 同じresidual formulationをretargeting後の動的実現可能性回復にも使用。

4. どうやって有効だと検証した?

- 25,454 single-hand trajectoriesを再構成(success 7.3%→59.3%)。 - 25 dual-hand tasksを再構成(16.0%→62.4%)。 - 各設定につき1つのshared policyを使用。 - 1つのhuman datasetを4つの形態の異なるrobot handsへ転移(+62.4 pp)。 - 実機で4つのcontact-rich bimanual tasksを実行。 - 実ロボット訓練データはゼロ。

5. 議論はある?

- 要旨からは不明。 - 限界、失敗事例、計算コスト、sim-to-real gap、汎化範囲についての議論は要旨に記載なし。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、human-to-robot retargeting、residual reinforcement learning、MANO hand model、motion-captureからのdemonstration learning、contact-rich manipulationが挙げられる。 - 同分野の定番として、Dexterous Manipulation from Human Demonstrations、Retargeting、Residual RL、Sim-to-Real Transferに関する研究を読むべき。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zihao Yang, Chengyuan Liu, Yu Zhou, Runze Lv, Tianyu Cui, Sheng Yi, Haohua Zhu, Irvine Lu, JieQ Sun

分類: cs.RO

原文アブストラクト

Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine whether a grasp succeeds are absent from every scalable source of human demonstrations. This paper builds on two observations. First, what survives the change from a human hand to a robot hand is the contact structure of a demonstration - which finger regions touch which object locations, and in what order - rather than its joint motion. Second, physical consistency need not be engineered per task: a single residual reinforcement learning (RL) policy, trained once across diverse demonstrations, can repair kinematic recordings into physically consistent, contact-annotated trajectories, and the same residual formulation restores dynamic feasibility after retargeting. These observations yield a three-stage pipeline that converts human motion-capture recordings into dexterous robot policies with no real-robot training data: physics refinement with a simulated MANO hand recovers contacts and forces, contact-anchored retargeting transfers the demonstrated contact structure through an objective independent of hand morphology, and residual policy learning adapts the result to robot actuation. The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.

関連論文

PR本紙発行元 EmplifAI