日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.07681

EigenDEXplore: 人間の事前知識を活用した構造的探索による巧みなマニピュレーション

EigenDEXplore: Structured Exploration for Dexterous Manipulation with Human Priors

シェア:XThreadsFacebookLINEはてブBluesky

人間の手の動作から得た固有ベクトルに沿ってノイズを加えることで、関節空間の探索を構造化し、巧みなマニピュレーションの強化学習性能を向上させる手法を提案。

詳しい要約

1. どんなもの?

- 器用な操作のための強化学習と軌道最適化において、人間の手の動作データから抽出した固有ベクトルに沿って摂動を加えることで、関節空間の探索を構造化する手法EigenDEXploreを提案。 - 行動空間を変更せず、独立した関節ノイズに相関のある探索を導入する。 - 多様な器用な手、把持、手内回転、接触の多い操作タスクで評価。

2. 先行研究と比べてどこがすごい?

- 先行研究では、人間の手データから学習した低次元の協調関節運動空間を用いて把持学習の探索空間を削減するが、一般的な操作に必要な表現力が制限される。 - 学習された行動空間と関節空間の行動を組み合わせて表現力を回復する手法もあるが、次元が増加し冗長性が生じる。 - EigenDEXploreは、人間の動作事前分布を行動表現の変更ではなく探索の構造化に用いることで、これらの問題を回避し、関節空間および学習行動空間のベースラインを一貫して上回る。

3. 技術・手法の肝は?

- 人間の手の動作データから固有ベクトルを抽出し、それに沿った摂動を独立した関節空間ノイズに追加することで、相関のある探索を実現。 - 行動空間は変更せず、探索のみを構造化する。 - これにより、協調的な動作の発見が容易になり、表現力を維持したまま探索効率を向上。

4. どうやって有効だと検証した?

- 複数の器用な手、把持、手内回転、接触の多い操作タスクにおいて、関節空間および学習行動空間のベースラインと比較。 - 非構造化強化学習と参照ガイド付き強化学習、軌道最適化、sim-to-real展開にわたって性能向上を確認。 - 報酬形成やカリキュラム設計が少ない設定で最大の効果を発揮。

5. 議論はある?

- 人間の動作事前分布は、行動表現を変更するよりも探索を構造化するために使用する方が効果的であることを実験的に示唆。 - 報酬形成やカリキュラム設計が少ない設定で特に有効。 - 限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されている先行研究:低次元の協調関節運動空間を用いた把持学習、学習行動空間と関節空間の行動を組み合わせた手法。 - 関連手法:強化学習、軌道最適化、sim-to-real転移。 - 同分野の定番:DAPG、Dexterous Manipulation with Deep Reinforcement Learningなど。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Harsh Gupta, Tyler Ga Wei Lum, Changhao Wang, Chuer Pan, C. Karen Liu, Jeannette Bohg, Shuran Song

分類: cs.RO, cs.AI

原文アブストラクト

Dexterous manipulation poses a challenging high-dimensional optimization problem, as useful behaviors require coordinated motion across many hand joints. In reinforcement learning (RL) and sampling-based trajectory optimization, exploration commonly relies on independent robot joint perturbations, making coordinated behaviors difficult to discover. Prior work reduces this search space for grasp learning using low-dimensional spaces of coordinated joint motions learned from human hand data, but this restricts the expressivity required for general manipulation. Some combine learned and joint-space actions to restore expressivity, but this increases dimensionality and introduces redundancy. We study these effects across diverse manipulation settings, varying action dimensionality, exploration strategy, and the source of human data. Our experiments suggest that human-motion priors are most effective when used to structure exploration rather than change the action representation. Motivated by this finding, we propose EigenDEXplore, which induces correlated exploration by adding perturbations along human-derived eigenvectors to independent joint-space noise, leaving the action space unchanged. Across multiple dexterous hands, EigenDEXplore consistently outperforms joint-space and learned action-space baselines in grasping, in-hand reorientation, and contact-rich manipulation. These gains span unstructured and reference-guided RL, trajectory optimization, and sim-to-real deployment, and are largest in settings with less reward shaping and curriculum design.

関連論文

PR本紙発行元 EmplifAI