Know-Your-Scene (KYS)-SLAM: ステレオ視覚SLAMにおける特徴マッチングのための階層的意味・運動事前分布
Know-Your-Scene (KYS)-SLAM: Hierarchical Semantic-Motion Priors for Feature Matching in Stereo Visual SLAM
ORB-SLAM3を拡張し、意味・パノプティック・運動の事前分布を階層的に統合して特徴マッチングの対応コストを連続的に調整するSLAM手法を提案。動的物体上の特徴を除外せず減衰させることで、バンドル調整に必要な幾何的対応を保持する。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Preeti Chatterjee, Jin Lu, Jin Sun, Suchendra M. Bhandarkar
分類: cs.CV, cs.RO
原文アブストラクト
Stereo visual SLAM systems built on local descriptors suffer from semantic ambiguity, instance-level confusion, and independently moving objects, each corrupting data association and accumulating as trajectory drift. Prevailing semantic and dynamic SLAM methods address this through binary feature rejection, sacrificing correspondence density for outlier suppression. We contend that contextual implausibility is better expressed as a graded quantity than an exclusion criterion. We present Know-Your-Scene (KYS)-SLAM, a modular extension of ORB-SLAM3 that supplants feature rejection with continuous correspondence modulation. The contribution is the reframing of contextual evidence as correspondence cost, applied within feature matching and leaving the geometric backend unmodified. Each keypoint is augmented with semantic, panoptic, and motion priors fused through a hierarchical compatibility formulation, in which semantic class and instance identity enforce structural plausibility while a zero-shot motion score down-weights features on independently moving objects. That score comes from a training-free module fitting a depth-aware ego-motion model to background optical flow and classifying panoptic segments via self-calibrating, coverage-aware thresholds, so only segments with sufficient motion evidence are penalized and static structure is left unpenalized. Penalizing correspondences rather than discarding them preserves the geometric support bundle adjustment depends on. Under one fixed configuration, no coefficient retuned per sequence or dataset, KYS-SLAM reduces per-sequence ATE RMSE by 17.4% on outdoor KITTI and 27.7% on indoor EuRoC across 21 stereo sequences with no regressions, and by 6.6% on dynamic subsets of KITTI Tracking and 17.8%, up to 31.2%, on Virtual KITTI 2 -- cross-domain transfer across outdoor driving, indoor flight, and synthetic imagery under one set of constants.
関連論文
- HuMemSLAM: 人間の記憶に着想を得た効率的な意味的場所認識による堅牢なVisual SLAMVisual SLAM
- 深層単眼Visual SLAMはスケール問題を克服したのか?ScaleMasterデータセットとベンチマークVisual SLAM
- NetVLADとFaissによるVisual SLAMのリアルタイムループ閉じ込め検出Visual SLAM
- 汎用3D事前情報を用いた動的環境向けVisual SLAMVisual SLAM
- TurboMap: 視覚SLAMのためのGPU高速化ローカルマッピングVisual SLAM
- RSV-SLAM: 屋内動的環境におけるリアルタイム意味論的Visual SLAMVisual SLAM