日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.11194

OmniDex: 多様な雑然シーンへの巧みな手把持のスケーリング

OmniDex: Scaling Dexterous Hand Grasping to Diverse Cluttered Scenes

シェア:XThreadsFacebookLINEはてブBluesky

260万以上の雑然シーンと4億の把持正解データからなる大規模ベンチマークを構築し、生成モデルの課題を克服するOmniDexモデルを提案した。

詳しい要約

1. どんなもの?

- 器用なロボットハンドによる grasping を、多様な clutter シーンへスケールさせる研究。 - 大規模データ不足がボトルネックである点に着目。 - 2.6M 超のシーンと 0.4B のシーン固有 grasp ground truth を持つ benchmark を構築。 - OmniDex モデルを提案し、生成モデルの課題を克服。 - 多様なシーン・視点・未見物体への汎化を目指す。

2. 先行研究と比べてどこがすごい?

- 従来は実世界データ収集が高コストで simulation が主流。 - しかし clutter シーンの大規模データは著しく不足していた。 - 既存の scene-level optimization は遅く、スケールしにくい。 - 本研究は seed-and-filter 戦略でこれを回避し、大規模 benchmark を実現。 - 既存の generative model が抱える grasp multimodality と last-millimeter precision errors に対処。 - post-optimization の遅延なしで robust grasping を達成。

3. 技術・手法の肝は?

- 高品質な 3D objects と supporting bases をキュレーション。 - 遅い scene-level optimization を回避する scalable seed-and-filter 戦略を提案。 - OmniDex モデルでは Soft Winner-Takes-All learning を採用。 - 訓練時に human-inspired physical constraints を組み合わせる。 - physics-driven ranking を利用。 - これにより multimodality と精度誤差を克服し、post-optimization 不要。

4. どうやって有効だと検証した?

- 2.6M 超のシーンと 0.4B の grasp ground truth からなる benchmark を構築。 - 多様な realistic layouts と rich semantic/geometric observations を含む。 - 実験結果で OmniDex が state-of-the-art 性能を達成。 - 多様な scenes, views, unseen objects への強い generalization を示す。

5. 議論はある?

- 要旨からは不明。 - ただし、clutter シーンでの grasping 学習におけるデータ不足の解決を主張。 - post-optimization の遅延を排除できる点を利点として提示。 - 限界や失敗事例、計算コストなどの議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として generative models for grasping、scene-level optimization、Soft Winner-Takes-All learning が挙げられる。 - 同分野の定番として dexterous grasping の simulation benchmark や clutter シーン grasping の研究を読むと良い。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Naiyu Fang, Zhongjin Luo, Yuxin Mo, Siyuan Huang, Jianbo Liu, Yufei Liu, Zheyuan Zhou, Chenkai Jin, Xiaogang Wang, Hongsheng Li

分類: cs.RO, cs.CV

原文アブストラクト

Dexterous grasping is the foundational primitive in embodied AI, demanding massive data to train robust models. As real-world data collection is expensive, simulation has become the mainstream paradigm. Yet, while cluttered scenes best reflect real-world applications, learning to grasp within them is bottlenecked by a critical scarcity of large-scale data. To resolve this, we curate high-quality 3D objects and supporting bases, proposing a scalable seed-and-filter strategy that bypasses sluggish scene-level optimization. This yields an unprecedented benchmark comprising over 2.6 million scenes and 0.4B scene-specific grasp ground truths, featuring diverse realistic layouts paired with rich semantic and geometric observations. Furthermore, we introduce the OmniDex model to overcome the grasp multimodality and last-millimeter precision errors plaguing current generative models. By coupling Soft Winner-Takes-All learning with human-inspired physical constraints during training, and utilizing physics-driven ranking, our approach achieves robust dexterous grasping without the latency of post-optimization. Experimental results show that OmniDex model achieves state-of-the-art performance and strong generalization across diverse scenes, views, and unseen objects.

関連論文

PR本紙発行元 EmplifAI