日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.12440

生成的ニューラルリターゲティングによる人間からロボットへの巧みな操作

Generative Neural Retargeting for Human-to-Robot Dexterous Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

フローマッチングモデルを用いて人間の動作をロボットの実行可能な軌道に変換する手法を提案し、MPCより少ないサンプルで高い成功率を達成した。

詳しい要約

1. どんなもの?

- 人間のデモンストレーションをロボットの器用な操作に転移するためのGenerative Neural Retargeting (GNR)を提案。 - 人間とロボットの身体性ギャップを埋めるため、流れマッチングモデルを用いて実現可能な軌道をサンプリングする。 - 大規模・長期間・ミリ精度の人間デモンストレーションを効率的にリターゲティングできる。 - 実世界からシミュレーションへのデータエンジンに適用し、223kデモと3.3k物体形状を含む接触力ラベル付きデータセットを生成。

2. 先行研究と比べてどこがすごい?

- 従来のIKは動力学を無視し、実行不可能な動作を生成することが多い。 - RLやサンプリングベースMPCは動的に実現可能な動作を生成するが、サンプル効率が悪くハイパーパラメータに敏感。 - RLは訓練が高コストで不安定、報酬設計が面倒。MPCは各軌道を独立に最適化し、サンプルコストがデータセットサイズとタスク難易度に応じて急増。 - GNRはMPCの8.5%のサンプル数で成功率56.20%を達成し、MPCの27.20%を大幅に上回る。

3. 技術・手法の肝は?

- 動的に実現可能な軌道はデモンストレーション間で共有される低次元多様体の近傍に集中するという仮説に基づく。 - リターゲティングを、人間の動きに条件付けられた多様体からのサンプリング問題に還元。 - 流れマッチングモデルを用いて実現可能な軌道をサンプリングする。 - 各デモンストレーションごとに新たな最適化問題を解くのではなく、学習された生成モデルから効率的にサンプリング。

4. どうやって有効だと検証した?

- GNRをMPCと比較し、MPCの8.5%のサンプル数で成功率56.20%を達成(MPCは27.20%)。 - 実世界からシミュレーションへのデータエンジンにGNRを適用し、223kデモと3.3k物体形状を含む大規模データセットを生成。 - 生成されたデータセットには密な接触力ラベルが含まれる。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:Inverse Kinematics (IK), Reinforcement Learning (RL), Model Predictive Control (MPC), flow matching model。 - 関連手法:Generative Neural Retargeting (GNR)自体。 - 同分野の定番:Human-to-Robot Dexterous Manipulation, Retargeting, Imitation Learning。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dechen Gao, Yue Yang, Ben Abbatematteo, Nathan Godwin, Pengcheng Wang, Roger Boldu, Steven Man, Zhiyang Dou, Chuan Qin, Sho Nakagome, Eric Whitmire

分類: cs.RO

原文アブストラクト

Human demonstrations are a scalable data source for learning dexterous manipulation, but the embodiment gap prevents human motion from being executed directly on robots. Inverse kinematics (IK) retargets human motion to robots efficiently but ignores dynamics, often producing infeasible motions. Reinforcement learning (RL) and sampling-based model predictive control (MPC) are commonly employed to yield dynamically feasible motions, but both are sample-inefficient and sensitive to hyperparameters. RL suffers from costly and unstable training and tedious reward engineering; MPC avoids policy optimization, yet retargets each trajectory in isolation, and solving one does not make the next easier. Sampling cost grows rapidly with dataset size and task difficulty. We hypothesize that dynamically feasible trajectories concentrate near a low-dimensional manifold shared across demonstrations, so that retargeting can be reduced to sampling from that manifold, conditioned on human motion, rather than solving a fresh optimization problem for every demonstration. We propose \textbf{Generative Neural Retargeting} (GNR), which uses a flow matching model to sample feasible trajectories. GNR outperforms MPC with only $8.5\%$ of the samples required by MPC, achieving a success rate of $56.20\%$ compared to $27.20\%$ for MPC. GNR can be used for scalable and efficient retargeting of large-scale, long-horizon, and millimeter precision human demonstrations: by applying GNR within a real-to-sim data engine, we produce a dexterous manipulation dataset with dense contact-force labels, spanning $223$k demonstrations and $3.3$k object geometries.

関連論文

PR本紙発行元 EmplifAI