日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.12470

Dex-One2Many: 単一の人間デモンストレーションから多様な巧みな操作を学習

Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration

シェア:XThreadsFacebookLINEはてブBluesky

人間の動画1本をシーングラフに抽象化し、強化学習の探索を導くことで、動画にない物体姿勢や把持にも汎化する多指ハンド操作方策を実機転移まで実現した。

詳しい要約

1. どんなもの?

- 単一の人間動画から汎化可能な器用な操作方策を学習する real-to-sim-to-real フレームワーク Dex-One2Many を提案。 - 動画を sequential scene graphs に抽象化し、RL の探索を導く。 - グラフは多様な reset states の生成制約と各段階の dense rewards を提供。 - シミュレーションのみで訓練し、実機の multi-fingered hand に zero-shot 転移。 - 5つの tool-use と manipulation タスクで評価。

2. 先行研究と比べてどこがすごい?

- 単一動画からの模倣は厳密な動作一致のため、未見の初期物体姿勢・目標姿勢・把持に汎化しにくい。 - RL は広く汎化可能だが、事前指導なしでは高次元探索が困難で多段階タスクに弱い。 - 提案法は scene graphs を生成制約と dense rewards に用い、探索を効率化しつつ広い汎化を両立。 - 結果として seen configurations で baseline を 6.5% 上回り、unseen scenarios では 71% の差に拡大。

3. 技術・手法の肝は?

- 人間動画を sequential scene graphs に抽象化。 - グラフを generative constraints として多様な reset states をサンプリング。 - 各段階に dense rewards を与え、探索を短く誘導。 - グラフは正確な姿勢ではなく関係を制約するため、動画外の物体姿勢や把持を含む reset states を生成。 - シミュレーションで訓練し、実機 multi-fingered hand に zero-shot 転移。

4. どうやって有効だと検証した?

- 5つの tool-use および manipulation タスクで評価。 - seen configurations で baseline を 6.5% 上回る。 - unseen scenarios では汎化により 71% の差に拡大。 - シミュレーション訓練から実機 multi-fingered hand への zero-shot 転移を実証。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、グラフ構築の自動化の程度などは記述されていない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、単一動画からの模倣学習、reinforcement learning (RL)、real-to-sim-to-real 転移、scene graphs を用いたタスク表現が挙げられる。 - 同分野の定番として、dexterous manipulation の模倣学習や RL ベースの手法を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jusuk Lee, Sungha Kim, Yeonsoo Park, Jonguk Cheon, Yoonkyo Jung, Yongjun You, H. Jin Kim, Jia-Bin Huang, Furong Huang, Youngseok Jang, Seungjae Lee

分類: cs.RO, cs.CV

原文アブストラクト

While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions. Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video. Alternatively, discovering a policy via reinforcement learning (RL) allows for broad generalization, but without prior guidance, it struggles with high-dimensional exploration in complex, multi-stage tasks. To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video. Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability. The graphs serve as generative constraints for sampling diverse reset states and provide dense rewards for each stage. Because the graphs constrain relations rather than exact poses, these reset states cover object poses and grasps beyond the video, while initializing each stage from them with dense rewards keeps exploration short and guided. Trained entirely in simulation, Dex-One2Many transfers zero-shot to a real multi-fingered hand. Across five tool-use and manipulation tasks, Dex-One2Many exceeds baselines by 6.5% in seen configurations, while its robust generalization widens this gap to 71% in unseen scenarios.

関連論文

PR本紙発行元 EmplifAI