日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.32626

World SLAM Model:SLAMとナビゲーションのための統合世界モデリング

World SLAM Model: Joint World Modeling for SLAM and Navigation

シェア:XThreadsFacebookLINEはてブBluesky

SLAMの逐次状態更新とバックエンド補正の仕組みをナビゲーションに直接組み込み、将来の視覚状態予測とカメラ運動・密な幾何推定を統合した世界モデルを提案。

著者: Minghui Qin, Yijun Yuan, Weicheng Zheng, Kenan Li, Weibang Wang, Chang Sun, Junhao Huang, Anmin Liu, Yicheng Yao, Hang Zhao

分類: cs.RO, cs.CV

原文アブストラクト

We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent memory and backend refinement of accumulated errors, to maintain a consistent world state during interaction. Given the current observation and a navigation goal, WSM predicts future visual states and jointly estimates their camera motion and dense geometry, grounding visual prediction in an evolving spatial world state. This spatial state is continuously updated as new observations arrive and provides the basis for action generation and closed-loop navigation. WSM is trained end-to-end with a joint navigation--SLAM objective, enabling downstream navigation to benefit directly from SLAM-style state maintenance and refinement while preserving accurate geometric estimation. Experiments demonstrate improved navigation performance together with strong SLAM accuracy, highlighting the potential of SLAM as an intrinsic mechanism for long-horizon world modeling and embodied interaction.

関連論文

PR本紙発行元 EmplifAI