日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.18718

混雑環境での把持のための視覚言語モデルによる校正済み確率的障害物推論

Calibrated Probabilistic Obstruction Reasoning with Vision-Language Models for Grasping in Clutter

シェア:XThreadsFacebookLINEはてブBluesky

VLM・深度・アモーダルマスクの手がかりを校正・統合し、障害物グラフ上の確率分布を推論して、把持・障害物除去・保留の判断を不確実性付きで行う枠組みを提案。

詳しい要約

1. どんなもの?

- 混雑環境で目標物体を把持するため、把持・障害物除去・保留を確率的に判断する枠組み - CPOR-Grasp という calibrated probabilistic obstruction-reasoning framework を提案 - VLM, depth, amodal-mask の手がかりを統合し、obstruction graph 上の不確実性を行動決定へ伝播 - 目標到達可能性や blocker 除去の尤度を計算し、certified decisions, adaptive stopping, principled deferral を可能にする

2. 先行研究と比べてどこがすごい?

- 既存手法は単一の obstruction graph や除去戦略にコミットし、代替解釈間の不確実性を無視 - 既存手法は miscalibrated な VLM 予測に依存し、pairwise obstruction 関係が jointly inconsistent になり得る - 既存の近似は破棄仮説が最終決定に与える影響の保証がない - CPOR-Grasp は calibration error を Gemini Robotics backbone で 0.1416 から 0.0185 に低減 - graph truncation で exact inference と 99.74% の決定で一致し、56 倍少ない graph で実現

3. 技術・手法の肝は?

- VLM, depth, amodal-mask cues を calibrate し fuse して obstruction probabilities を推定 - valid obstruction graphs 上の分布を誘導し、それらについて marginalize して目標到達可能性や blocker 除去の尤度を計算 - 推論を tractable にするため最高確率の graph のみを保持 - 破棄された確率質量に対する total-variation bound を導出し、certified decisions, adaptive stopping, principled deferral を実現

4. どうやって有効だと検証した?

- synthetic および real UNOBench scenes で評価 - state-of-the-art baselines を上回る性能 - Gemini Robotics backbone で calibration error が 0.1416 から 0.0185 に減少 - graph truncation が exact inference と 99.74% の決定で一致し、56 倍少ない graph を使用 - 実世界実験で平均成功率 77.8% を達成し SOTA baselines を上回る

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- UNOBench 関連の研究 - Gemini Robotics backbone を用いた VLM 研究 - obstruction graph や clutter grasping に関する先行研究 - calibrated probabilistic reasoning や total-variation bound を用いた近似推論の研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Thanh-Tuan Tran, Ngoc-Chien Chu, Thanh Nguyen Canh, Nak Young Chong, Nguyen-Viet Ha, Xiem HoangVan

分類: cs.RO

原文アブストラクト

Retrieving a target from clutter requires deciding whether to grasp the target, remove a blocker, or defer. Existing methods typically commit to a single obstruction graph or removal strategy, ignoring uncertainty across alternative scene interpretations. They also rely on miscalibrated vision-language model (VLM) predictions and can produce pairwise obstruction relations that are jointly inconsistent. Moreover, current approximations provide no guarantees about the impact of discarded hypotheses on the final decision. We propose CPOR-Grasp, a calibrated probabilistic obstruction-reasoning framework that propagates uncertainty from pairwise evidence to action decisions. CPOR-Grasp calibrates and fuses VLM, depth, and amodal-mask cues to estimate obstruction probabilities, induces a distribution over valid obstruction graphs, and marginalizes over these graphs to compute the likelihood that the target is accessible or that a given blocker should be removed. To make inference tractable, it retains only the highest-probability graphs and derives a total-variation bound on the discarded probability mass, enabling certified decisions, adaptive stopping, and principled deferral. On synthetic and real UNOBench scenes, CPOR-Grasp outperforms state-of-the-art baselines. Calibration error decreases from 0.1416 to 0.0185 on the Gemini Robotics backbone, while graph truncation matches exact inference on 99.74\% of decisions using 56 times fewer graphs. In real-world experiments, CPOR-Grasp achieves a 77.8\% average success rate, surpassing SOTA baselines.

関連論文

PR本紙発行元 EmplifAI