EquiGQNet: 共有等変点群符号化による高速な把握品質評価
EquiGQNet: Fast Grasp Quality Evaluation via Shared Equivariant Point Cloud Encoding
6自由度の把握候補を高速かつ正確に評価するため、SO(3)等変な「一度符号化してから回転」方式と中間レベル行動融合を組み合わせた把握品質評価器EquiGQNetを提案した。シミュレーションと実機で従来法と同等以上の性能を保ちつつ、計画時間を6.9倍高速化した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Sungwon Seo, Jaeseog Won, Jiyou Shin, Youngjin Seo, Hyunjun Kim, Seokmin Yoon, Tuan Luong, Hyungpil Moon
分類: cs.RO
原文アブストラクト
Planning six-degree-of-freedom (6-DoF) grasps for unseen objects in cluttered tabletop scenes from a single-view depth image requires accurate and efficient evaluation of diverse grasp candidates. Existing early-fusion methods capture local object geometry relative to each grasp candidate but repeatedly encode the scene, whereas late-fusion methods reuse a shared scene representation but may lose this grasp-relative local geometry. We propose EquiGQNet, an efficient 6-DoF grasp quality evaluator that combines the strengths of both approaches. For grasp orientation, EquiGQNet replaces the early-fusion operation of rotating and re-encoding the point cloud for each grasp candidate with an SO(3)-equivariant encode-once-then-rotate scheme, yielding grasp-aligned geometric features from a shared scene encoding. For grasp translation, Mid-level Action Fusion (MAF) injects the grasp position into intermediate features before global aggregation, retaining local geometry relative to each candidate. We evaluate EquiGQNet in two grasp planning pipelines: Cross-Entropy Method (CEM)-based continuous grasp search and candidate ranking with a pretrained generative planner. In simulation, EquiGQNet achieves grasping performance comparable to the early-fusion baseline and substantially outperforms late fusion on objects with complex geometry and limited graspable regions, while reducing CEM planning time from 3.31s to 0.48s, a 6.9x speedup over early fusion. In real-world household-object decluttering, EquiGQNet achieves a 95.2% grasp success rate and 230 picks per hour, versus 153 and 170 for early- and late-fusion baselines. Code is available at https://equigqnet.github.io/.
関連論文
- FOCIポリシー:関係的操作ポリシーのためのオブジェクト中心相互作用に焦点を当てるマニピュレーション
- 3DWay: 3D一貫性ウェイポイントによるロボット操作の一般化マニピュレーション
- CASD: チャンク整合セマンティック蒸留による多段階ロボット操作マニピュレーション
- WM-Craftnet: 汎用かつ堅牢な器用な手内操作のための世界共感覚モデルマニピュレーション
- Dex-X: シミュレーションによるインタラクションを用いた人間のビデオからの視覚-触覚器用操作学習マニピュレーション
- VLA-Corrector: 段階認識可能な観測状態理解に基づくプロンプト駆動型閉ループ回復フレームワークマニピュレーション