日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
SLAMarXiv:2602.15899

SceneVGGT:VGGTベースのオンライン3DセマンティックSLAMによる屋内シーン理解とナビゲーション

SceneVGGT: VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation

シェア:XThreadsFacebookLINEはてブBluesky

VGGTを基盤にスライディングウィンドウで長尺動画を処理し、2Dインスタンスマスクを3Dオブジェクトに持ち上げて時系列一貫性を保つオンライン3DセマンティックSLAMを構築。推定床面に物体位置を投影し、音声フィードバック付き支援ナビゲーションを実現した。

著者: Anna Gelencsér-Horváth, Gergely Dinya, Dorka Boglárka Erős, Péter Halász, Islam Muhammad Muqsit, Kristóf Karacs

分類: cs.RO, eess.IV

原文アブストラクト

We present SceneVGGT, a spatio-temporal 3D scene understanding framework that combines SLAM with semantic mapping for autonomous and assistive navigation. Built on VGGT, our method scales to long video streams via a sliding-window pipeline. We align local submaps using camera-pose transformations, enabling memory- and speed-efficient mapping while preserving geometric consistency. Semantics are lifted from 2D instance masks to 3D objects using the VGGT tracking head, maintaining temporally coherent identities for change detection. As a proof of concept, object locations are projected onto an estimated floor plane for assistive navigation. The pipeline's GPU memory usage remains under 17 GB, irrespectively of the length of the input sequence and achieves competitive point-cloud performance on the ScanNet++ benchmark. Overall, SceneVGGT ensures robust semantic identification and is fast enough to support interactive assistive navigation with audio feedback.

関連論文