日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.18117

OpenDexGrasp: オープン語彙のタスク指向巧み把持

OpenDexGrasp: Open-vocabulary Task-Oriented Dexterous Grasping

シェア:XThreadsFacebookLINEはてブBluesky

自由形式の言語指示から機能意図を推論し、多視点視覚と物体形状に基づいて高自由度の巧みな把持を生成する統合フレームワークを提案。

詳しい要約

1. どんなもの?

- 本研究は、open-vocabularyなtask-oriented dexterous grasp生成を扱う。 - 自由形式の言語から機能的意図を推論し、multi-view視覚観測と物体形状に基づき、実行可能な高自由度graspを生成する。 - OpenDexGraspは、この設定のための統合データ・生成モデリングフレームワーク。 - OpenDexVerseは、Coverage-to-Alignment (C2A) Recipeで整理されたdual-source supervisionを提供する。 - OpenDex-Scaleは自動grasp合成とvision-language注釈で大規模な意味・幾何カバレッジを提供。 - OpenDex-Alignは人間のteleoperationとカテゴリレベル転移で高品質なembodied alignmentを提供。

2. 先行研究と比べてどこがすごい?

- 従来のdexterous grasp合成は安定で物理的に妥当な手姿勢生成が進歩したが、タスクが示す機能を保持するgraspは不十分。 - 本研究はopen-vocabularyなtask-oriented設定を扱い、言語から機能意図を推論する点が新しい。 - 別個のaffordance-to-pose推論段階を必要とせず、task-consistentなdexterous graspを直接生成する。 - シミュレーションと実機で、機能的整合性、物理的実現可能性、未見カテゴリへの汎化、実世界実行成功率の改善を示す。

3. 技術・手法の肝は?

- OpenDexGraspは、open-vocabularyなvision-language文脈とdexterous action生成を結合する共有perception-action latent representationを学習する。 - Affordance groundingとgrasp generationがこのlatent space上で相補的なsupervisionを提供する。 - これにより、別個のaffordance-to-pose推論段階なしで直接task-consistentなdexterous graspを生成する。 - データ面では、OpenDexVerseがC2A Recipeに基づくdual-source supervisionを提供する。 - OpenDex-Scaleは自動grasp合成とvision-language注釈、OpenDex-Alignは人間のteleoperationとカテゴリレベル転移によるembodied alignmentを供給する。

4. どうやって有効だと検証した?

- 広範なシミュレーションと実ロボット実験を実施。 - 機能的整合性、物理的実現可能性、未見カテゴリへの汎化、実世界実行成功率の改善を実証。 - 詳細と動画は https://opendexgrasp.github.io/ で公開。

5. 議論はある?

- 要旨からは、限界や失敗事例、計算コスト、データ収集のバイアスなどに関する具体的な議論は不明。 - 提案手法の有効性と汎化性が強調されているが、定量的な比較やアブレーションの詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別の先行研究は明示されていない。 - 関連手法として、dexterous grasp synthesis、open-vocabulary grasping、affordance grounding、vision-language-action models、human teleoperationによるデータ収集などが挙げられる。 - 同分野の定番として、DexGraspNet、GraspNet、CLIP、RT-1/RT-2などの一般名が考えられるが、要旨に直接の言及はない。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiyao Zhang, Junhan Wang, Tianyu Wang, Zeyuan Chen, Anthony Bolten, Yitong Peng, Hao Dong

分類: cs.RO

原文アブストラクト

Dexterous grasp synthesis has advanced rapidly in generating stable and physically plausible hand poses, but real-world manipulation requires grasps that preserve the function implied by the task. We study open-vocabulary task-oriented dexterous grasp generation, where a robot must infer functional intent from free-form language, ground it in multi-view visual observations and object geometry, and generate an executable high-degree-of-freedom grasp. We present OpenDexGrasp, a unified data and generative modeling framework for this setting. OpenDexVerse provides dual-source supervision organized by the Coverage-to-Alignment (C2A) Recipe: OpenDex-Scale offers large-scale semantic and geometric coverage through automatic grasp synthesis and vision-language annotation, while OpenDex-Align supplies high-quality embodied alignment through human teleoperation and category-level transfer. OpenDexGrasp learns a shared perception-action latent representation that couples open-vocabulary vision-language context with dexterous action generation. Affordance grounding and grasp generation provide complementary supervision over this latent space, enabling direct generation of task-consistent dexterous grasps without a separate affordance-to-pose inference stage. Extensive simulation and real-robot experiments demonstrate improved functional alignment, physical feasibility, generalization to unseen categories, and real-world execution success. Additional details and videos are available at https://opendexgrasp.github.io/.

関連論文

PR本紙発行元 EmplifAI