日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
巧みな操作/VLA/身体間転移arXiv:2608.14028v1

AdvDex: 関節整合アクションと敵対的学習による人間のデモからの巧みな操作学習

AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning

シェア:XThreadsFacebookLINEはてブBluesky

人間とロボットのデモを統合したVLAフレームワークを提案し、関節整合アクション空間と敵対的学習により異なる身体間の汎化を実現した。

詳しい要約

1. どんなもの?

AdvDexは、人間とロボットのデモンストレーションから器用な操作を学習するための統一的なVision-Language-Actionフレームワークである。大規模マルチモーダルデータセットOmniShare、正準行動表現であるJoint-Aligned Action Space (JAAS)、およびドメイン逆学習を導入し、異なる身体性間の一般化を実現する。

2. 先行研究と比べてどこがすごい?

従来の手法は、ロボットのテレオペレーションによる高コストなデータ収集に依存し、異なる身体性間の行動空間の違いや視覚表現の身体性固有の情報への絡まりに対処できなかった。AdvDexは、人間のデモンストレーションから高品質なキネマティック監視と触覚測定を提供するOmniShareを導入し、JAASにより人間の手、器用なロボットハンド、パラレルグリッパーを機能的に整合させる。さらに、ドメイン逆学習により視覚表現から身体性固有の情報を除去し、クロス身体性一般化を向上させる。

3. 技術・手法の肝は?

手法の肝は3点。(1) OmniShare: 人間の操作デモンストレーションを大規模に収集したマルチモーダルデータセットで、キネマティック監視と触覚測定を提供し、ロボットのテレオペレーションへの依存を低減。(2) JAAS: SE(3)手首姿勢と15指関節からなる正準行動表現で、人間の手、器用なロボットハンド、パラレルグリッパーを機能的に整合させる。(3) ドメイン逆学習: 視覚表現から身体性固有の情報を除去し、タスク関連の視覚手がかりを保持する。

4. どうやって有効だと検証した?

ハンドアクション予測と実世界の器用な操作タスクにおいて、ベースラインと比較して一貫した改善を示した。ゼロショットの人間からロボットへのスキル転送、未見の物体や環境への一般化、データ効率的な少数ショット適応を実証した。

5. 議論はある?

要旨からは、提案手法の限界や潜在的な欠点についての議論は不明である。ただし、ドメイン逆学習が視覚表現から身体性情報を除去する際に、タスク関連情報も失う可能性や、JAASがすべての身体性に適合するかどうかなどの課題が考えられるが、要旨には明記されていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、DexMV、DexPBT、RoboTurkなどのロボットデモンストレーション学習、およびVision-Language-Actionモデル(例:RT-2、Octo)が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhiyue Zhao, Jingyi Wu, Hairuo Liu, Mingyu Liu, Liyang Li, Hengdi Zhang, Tong He, Zhengxue Cheng

分類: cs.RO, cs.AI

原文アブストラクト

Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spaces vary across embodiments. Policies trained on heterogeneous data can also entangle task-relevant visual cues with embodiment-specific appearance, limiting cross-embodiment generalization. We present AdvDex, a unified Vision-Language-Action framework for learning dexterous manipulation from human and robot demonstrations. First, we introduce OmniShare, a large-scale multimodal dataset of human manipulation demonstrations that provides high-quality kinematic supervision and tactile measurements while reducing reliance on robot teleoperation. Second, we propose the Joint-Aligned Action Space (JAAS), a canonical action representation comprising an $\mathrm{SE}(3)$ wrist pose and 15 finger joints, thereby functionally aligning human hands, dexterous robot hands, and parallel grippers. Finally, we use domain-adversarial learning to reduce embodiment-specific information in the learned visual representation. Experiments on hand-action prediction and real-world dexterous manipulation show consistent improvements over baselines, effective zero-shot human-to-robot skill transfer, generalization to unseen objects and environments, and data-efficient few-shot adaptation.