混雑環境での把持のための視覚言語モデルによる校正済み確率的障害物推論
Calibrated Probabilistic Obstruction Reasoning with Vision-Language Models for Grasping in Clutter
VLM・深度・アモーダルマスクの手がかりを校正・統合し、障害物グラフ上の確率分布を推論して、把持・障害物除去・保留の判断を不確実性付きで行う枠組みを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Thanh-Tuan Tran, Ngoc-Chien Chu, Thanh Nguyen Canh, Nak Young Chong, Nguyen-Viet Ha, Xiem HoangVan
分類: cs.RO
原文アブストラクト
Retrieving a target from clutter requires deciding whether to grasp the target, remove a blocker, or defer. Existing methods typically commit to a single obstruction graph or removal strategy, ignoring uncertainty across alternative scene interpretations. They also rely on miscalibrated vision-language model (VLM) predictions and can produce pairwise obstruction relations that are jointly inconsistent. Moreover, current approximations provide no guarantees about the impact of discarded hypotheses on the final decision. We propose CPOR-Grasp, a calibrated probabilistic obstruction-reasoning framework that propagates uncertainty from pairwise evidence to action decisions. CPOR-Grasp calibrates and fuses VLM, depth, and amodal-mask cues to estimate obstruction probabilities, induces a distribution over valid obstruction graphs, and marginalizes over these graphs to compute the likelihood that the target is accessible or that a given blocker should be removed. To make inference tractable, it retains only the highest-probability graphs and derives a total-variation bound on the discarded probability mass, enabling certified decisions, adaptive stopping, and principled deferral. On synthetic and real UNOBench scenes, CPOR-Grasp outperforms state-of-the-art baselines. Calibration error decreases from 0.1416 to 0.0185 on the Gemini Robotics backbone, while graph truncation matches exact inference on 99.74\% of decisions using 56 times fewer graphs. In real-world experiments, CPOR-Grasp achieves a 77.8\% average success rate, surpassing SOTA baselines.
関連論文
- SeeQ: 長期的ロボットマニピュレーションのための汎用価値関数の学習マニピュレーション
- グリッパを考慮した不規則物体の自動高密度パッキングマニピュレーション
- 事前学習から熟達へ:最小限の人的介入で長期的マニピュレーションを実現する実世界サブタスクRLマニピュレーション
- ForceTwin: 計測された人間の操作からロボットマニピュレーションのための物理情報デジタルツインを構築マニピュレーション
- 並列シミュレーションにおけるロボットマニピュレーションのための視覚言語報酬学習のスケーリングマニピュレーション
- 細粒度物体操作に向けて:SAM3誘導視覚運動ポリシーと持続的メモリ学習および集中視覚条件付けマニピュレーション