日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.01849

FlashDexRetarget: 多動作リターゲティングによる器用操作データ生成の高速化

FlashDexRetarget: Accelerating Dexterous Manipulation Data Generation through Multi-Motion Retargeting

シェア:XThreadsFacebookLINEはてブBluesky

人間の手と物体のデモンストレーションをロボットハンドに転送する強化学習ベースのリターゲティング手法を提案し、50動作で90%の成功率を達成しつつ、学習計算量を最大100倍削減した。

詳しい要約

1. どんなもの?

- 人間の手と物体のデモンストレーションを、器用なロボットハンドの操作データへ変換する手法。 - 物理的に実行可能な retargeting を目的とした RL ベースのフレームワーク。 - 単一物体と二物体の相互作用を含む 50 モーションのベンチマークで評価。 - XHand と Sharpa Wave Hand の両方で一貫した性能向上を確認。 - 200、500、1,000 モーションへスケールしても安定し、大規模化で効率が向上。

2. 先行研究と比べてどこがすごい?

- 既存の physics-based 手法は retargeting 成功率、モーション固有の学習効率、またはその両方に制限があった。 - 提案手法は 50 モーションで 90% の成功率を達成。 - 評価した sampling-based ベースラインの約 2.5 倍の成功率。 - 評価した RL-based ベースラインより最大 100 倍少ない学習計算量で実現。 - 大規模モーション数でも安定し、学習セット拡大に伴い効率的に成功モーションを生成。

3. 技術・手法の肝は?

- RL ベースの高成功率・高効率な dexterous motion retargeting フレームワーク。 - 物体 point-cloud 観測、hand-object 距離特徴、future trajectory encodings を組み合わせる。 - 物体運動と参照 hand-object 関係を監督する補完的報酬を併用。 - 左右の手それぞれに separate actor critic networks を採用。 - off-policy アルゴリズム FlashSAC を dexterous motion tracking に適応。

4. どうやって有効だと検証した?

- 単一物体・二物体相互作用を含む 50 モーションのベンチマークで評価。 - 成功率 90% を達成し、sampling-based ベースラインの約 2.5 倍。 - RL-based ベースラインより最大 100 倍少ない学習計算量。 - XHand と Sharpa Wave Hand で一貫した性能向上を確認。 - component-wise ablations で設計選択の寄与を検証。 - 200、500、1,000 モーションでスケール安定性を確認。 - 実世界で取得したデモンストレーションの qualitative replay 結果も提示。

5. 議論はある?

- 要旨からは不明。 - ただし component-wise ablations により設計選択の寄与を検討している。 - 大規模モーション数での安定性と効率向上を報告。 - 実世界キャプチャのデモへの適用可能性を定性的に示す。

6. 次に読むべき論文は?

- 要旨で参照・比較されている sampling-based baselines および RL-based baselines。 - 提案手法が適応した off-policy アルゴリズム FlashSAC。 - 評価に用いられた XHand と Sharpa Wave Hand に関する研究。 - 関連する dexterous manipulation retargeting や physics-based retargeting の先行研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kyungmin Lee, Sibeen Kim, Dongyoon Hwang, Yoonsang Oh, Donghu Kim, Youngdo Lee, I Made Aswin Nahrendra, Jaegul Choo, Hojoon Lee

分類: cs.RO

原文アブストラクト

Human hand-object demonstrations offer a reusable source of dexterous robot manipulation data, but transferring them across embodiments requires physically feasible retargeting. Existing physics-based approaches face limitations in retargeting success, motion-specific training efficiency, or both. To address these limitations, we introduce FlashDexRetarget, an RL-based framework for high-success, efficient dexterous motion retargeting. To make the demonstrated interaction easier to learn, we combine object point-cloud observations, hand-object distance features, and future trajectory encodings with complementary rewards that supervise object motion and reference hand-object relationships. To further accelerate learning, we employ separate left- and right-hand actor critic networks and adapt the off-policy algorithm, FlashSAC to dexterous motion tracking. On a benchmark of 50 motions spanning single-object and two-object interactions, FlashDexRetarget achieves a 90% success rate, approximately 2.5x that of the evaluated sampling-based baselines, while requiring up to 100x less training compute than the evaluated RL-based baselines. Evaluations on both XHand and Sharpa Wave Hand show consistent gains, and component-wise ablations examine the contributions of our design choices. Beyond the 50-motion benchmark, experiments with 200, 500, and 1,000 motions demonstrate that our method remains stable at larger scales and produces successful retargeted motions more efficiently as the training set grows. Qualitative replay results using real-world-captured demonstrations further illustrate the applicability of our framework to recorded human manipulation. Videos and code are available at https://davian-robotics.github.io/FlashDexRetarget/

関連論文

PR本紙発行元 EmplifAI