日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.02493v1

MS-MEM: 不確実性・外乱を考慮した行動選択によるマルチスキル操作強化マッピング

MS-MEM: Multi-Skill Manipulation-Enhanced Mapping via Uncertainty- and Disturbance-Aware Action Selection

シェア:XThreadsFacebookLINEはてブBluesky

棚などの閉鎖的で散らかった環境でのシーン理解を高めるため、能動視点選択・物体押し・把持を統合し、不確実性を考慮したマッピングと外乱制約付き行動選択を提案した。

詳しい要約

1. どんなもの?

MS-MEMは、棚などの閉鎖的で雑然とした環境におけるサービスロボットのシーン理解を向上させるための、マルチスキル操作強化マッピングフレームワークである。能動的視点選択、物体押し、把持を統合し、不確実性を考慮したマッピングを実現する。シーンレベルのメトリック・セマンティックな証拠に基づく信念推定器と、不確実性を考慮した把持表現を組み合わせる。

2. 先行研究と比べてどこがすごい?

従来の研究では、単一のスキル(視点選択、押し、把持)を用いた能動的マッピングが多く、複数スキルの相乗効果を考慮したものは少ない。また、操作によるシーンの過度な変化(collateral disturbance)を考慮せず、マッピング精度のみを最適化していた。MS-MEMは、複数スキルを統合し、さらにシーン信念の確信領域への過度な変化を抑制する制約を導入することで、マッピング精度とシーン変化の低減を両立している点が新しい。

3. 技術・手法の肝は?

手法の核は、(1) シーンレベルのメトリック・セマンティックな証拠に基づく信念推定器を用いた不確実性の定量化、(2) 新しいfull-evidential grasp estimatorによる把持可能性と方向の不確実性のモデル化、(3) 統一的なアクション選択パイプラインにおける共通の情報利得基準による候補アクションの評価、(4) 操作アクションに対するcollateral disturbance constraint (CDC) の導入である。CDCは、シーン信念の確信領域への過度な変更を抑制し、不確実性低減とシーン変化のバランスを取る。

4. どうやって有効だと検証した?

実験では、単一スキルベースラインと、シーン変化を無視した制約なしベースラインと比較し、MS-MEMがより高いマッピング精度を達成しつつ、シーン変化を大幅に低減することを示した。これにより、能動的視点選択、押し、把持の相乗効果を実証した。

5. 議論はある?

要旨からは、提案手法の限界や特定の環境条件に関する議論は不明である。また、CDCの重み設定や、異なる環境での汎用性、計算コストなどについての詳細な議論は要旨には含まれていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていないが、同分野の定番として、能動的視点選択(Active Viewpoint Selection)、物体押し操作(Pushing)、把持計画(Grasp Planning)、証拠理論(Evidential Theory)に基づく不確実性推定、メトリック・セマンティックマッピング(Metric-Semantic Mapping)に関する論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yitian Shi, Jesper Mücke, Nils Dengler, Sicong Pan, Rania Rayyes, Maren Bennewitz

分類: cs.RO

原文アブストラクト

Accurate scene understanding in confined, cluttered spaces such as shelves is essential for service robots, as many everyday tasks require them to locate and retrieve objects reliably. Yet, it remains challenging due to severe occlusions, restricted accessibility, and the need to avoid excessive scene changes. In this paper, we propose Multi-Skill Manipulation-Enhanced Mapping (MS-MEM), an evidential framework for uncertainty-aware mapping that integrates active viewpoint selection, object pushing, and grasping. MS-MEM combines scene-level metric-semantic evidential belief estimators with an uncertainty-aware grasp representation. This representation is learned using a novel full-evidential grasp estimator that models both grasp affordance and orientation uncertainty. In our framework, candidate perception and manipulation actions are evaluated within a unified action selection pipeline using a common information gain criterion. For manipulation actions, we further introduce a collateral disturbance constraint (CDC) that discourages excessive changes to confident regions of the scene belief. This enables MS-MEM to select actions that effectively reduce map uncertainty while limiting collateral scene changes. Experimental results show that, compared with single-skill and unconstrained baselines that ignore scene disturbance, MS-MEM achieves higher mapping accuracy while substantially reducing scene disturbance, highlighting the synergistic effects of active viewpoint selection, push, and grasp actions.

関連論文