DRS-VPT: ビジョンポイントトランスフォーマーによるスキャンへの直接再ローカライゼーション
DRS-VPT: Directly Relocalizing in a Scan with Vision Point Transformers
クエリ画像と参照3D点群からスキャンの姿勢と点マップを予測し、カメラ-LiDARキャリブレーションや屋内再ローカライゼーションを統合的に行うフィードフォワードトランスフォーマーを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Lanke Frank Tarimo Fu, Maurice Fallon
分類: cs.CV, cs.RO
原文アブストラクト
We present DRS-VPT, a feed-forward transformer architecture for foundational image-to-scan registration. Given query images and a reference 3D point cloud, the model predicts the scan pose and point map alongside the poses and point maps of each camera, all expressed in the first camera's frame. It additionally predicts a coarse-to- fine pyramid of per-point and per-pixel features for direct reprojective alignment of the scan to the first image. This formulation unifies downstream tasks such as camera-LiDAR calibration in autonomous driving and indoor camera-to-map relocalization. A single DRS-VPT model achieves state-of-the-art performance for image-to-LiDAR registration in autonomous driving, competitive indoor relocalization without training map-specific weights, and strong zero-shot transfer to unseen environments. We also show qualitatively that the model learns complex scan-to-image projection properties such as occlusion of back-facing points.