不確実性下でのマルチモーダル指示グラウンディングによるマニピュレーション計画
MIGU: Multimodal Instruction Grounding under Uncertainty for Manipulation Planning
言語とジェスチャーの不確実な手がかりを統合し、VLMの意味情報と3D幾何学的尤度をベイズ的に融合して対象物体を推定、必要に応じて確認質問も行うマニピュレーション計画フレームワークMIGUを提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Mingke Lu, Anxing Xiao, David Hsu
分類: cs.RO
原文アブストラクト
Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io
関連論文
- DexTacWAM: 巧みな操作のための視触覚ワールドアクションモデルマニピュレーション
- 平面果樹園における視覚運動ロボット剪定のためのハイブリッド強化学習マニピュレーション
- CAST: 衝突を考慮した建設ロボットによる同時軌道推定と計画マニピュレーション
- ロボット構成空間における異種制約のための実行可能性距離場マニピュレーション
- InsertAnything: シミュレーションから現実への汎化可能な接触リッチ精密挿入マニピュレーション
- フレーズ単位のロボット古琴演奏:両腕動作計画と音触覚インタラクション監視マニピュレーション