日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.04438

RAGrasp: 幾何・意味テンプレート検索と把持転移

RAGrasp: Geometry-Semantic Template Retrieval and Grasp Transfer

シェア:XThreadsFacebookLINEはてブBluesky

展開環境で収集した少数のRGB-D把持テンプレートを検索し、SAM2とDINOv2特徴で対象物を切り出して把持を転移する、再学習不要の平面パラレルジョー把持パイプライン。

詳しい要約

1. どんなもの?

- RAGraspは、planar parallel-jaw graspingのためのretrieval-augmented pipeline。 - 展開環境で収集した少数のgrasp-annotated RGB-D templatesからgraspを生成。 - 新しい展開先でもend-to-end retrainingが不要。 - テンプレートメモリは展開カメラ・ロボット・グリッパで収集した観測から構築。

2. 先行研究と比べてどこがすごい?

- 従来は大規模公開データセットや合成graspデータセットで訓練されたtask-specific predictorsが主流。 - RAGraspは新展開ごとのend-to-end retrainingを必要としない。 - テンプレートメモリを展開環境のセンシング・embodiment条件に合わせて構築。 - 限られたローカルアノテーションで展開固有のgrasp適応を実現。

3. 技術・手法の肝は?

- self-supervised DINOv2 visual featuresとappearance・depth cuesでSAM2をpromptし、query objectを分離。 - two-stage geometry-semantic retrieval cascadeでテンプレートを選択。 - confidence gateが2つのgrasp-transfer estimatorsの一方を選択。 - 転送graspをmask-supportとsilhouette-contact constraintsでrefine。 - 最後にcalibrated 2D-to-3D conversionを実施。

4. どうやって有効だと検証した?

- 実世界トライアルで評価。 - seen objectsで20/20のgrasp成功。 - unseen objectsで19/20のgrasp成功。 - 評価設定内で、限られたローカルアノテーションからの展開固有grasp適応と、テストした視点・照明変化への耐性を示す。

5. 議論はある?

- 評価設定内での結果であり、展開固有のgrasp適応と視点・照明変化への耐性を示す。 - 限られたローカルアノテーションで適応可能。 - 一般化可能性や他の条件への拡張性は要旨からは不明。

6. 次に読むべき論文は?

- DINOv2 - SAM2 (Segment Anything Model 2) - retrieval-augmented grasping - planar parallel-jaw grasping - grasp transfer

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shenzhe Zhu, Chengxiao He, Jan Harder

分類: cs.AI, cs.RO

原文アブストラクト

We present RAGrasp, a retrieval-augmented pipeline for planar parallel-jaw grasping from a compact set of locally collected, grasp-annotated RGB-D (color and depth) templates. Unlike task-specific predictors trained primarily on large public or synthetic grasp datasets, RAGrasp requires no end-to-end retraining for a new deployment.Its template memory is constructed from observations collected with the deployment camera, robot, and gripper in the target workspace, thereby aligning stored examples with the local sensing and embodiment conditions. The system uses self-supervised DINOv2 visual fea- tures together with appearance and depth cues to prompt the Segment Anything Model 2 (SAM2), which isolates the query object. A two-stage geometry-semantic retrieval cascade then selects a template, and a confidence gate chooses one of two grasp- transfer estimators. The transferred grasp is refined using mask- support and silhouette-contact constraints before calibrated 2D- to-3D conversion. In real-world trials, RAGrasp achieves 20/20 successful grasps on seen objects and 19/20 on unseen objects. Within the evaluated setting, the results demonstrate deployment- specific grasp adaptation from limited local annotation and tolerance to the tested viewpoint and illumination changes.

関連論文

PR本紙発行元 EmplifAI