日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.29835

検索して特定:大規模言語モデルとLiDAR幾何学を橋渡しする空間グラウンディング

Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

シェア:XThreadsFacebookLINEはてブBluesky

LiDAR点群と言語モデルを組み合わせ、複雑な空間関係の質問に対して対象物の座標を高精度に特定する手法を提案。

詳しい要約

1. どんなもの?

- LiDARの幾何情報とLLMの言語事前知識を組み合わせ、複雑な空間質問に答えて対象物をLiDAR幾何でgroundingする研究。 - データセットSpatialLiDAR-QAを導入。単一・多段階の関係groundingと補完的な空間理解タスクを含む。 - モデルSpatialLiDAR-LMを提案。LiDAR点特徴をLLMに整合させ、言語条件付きの位置認識proposal retrievalと局所点refinementで対象座標をgroundingする。

2. 先行研究と比べてどこがすごい?

- 従来のLiDAR知覚は個別物体の認識・定位は可能だが、空間関係を構成して意図対象をgroundingする質問には不十分。 - 既存のLiDAR–languageモデルやmulti-camera VLMと比較し、精密な座標予測タスクで大幅改善。 - テキスト言語デコードではなく局所LiDAR幾何から直接座標を導出する点が異なる。

3. 技術・手法の肝は?

- LiDAR点特徴をLLMに整合させる。 - 言語条件付き・位置認識のproposal retrievalを実行。 - 局所点refinementにより対象座標をLiDAR幾何から直接導出。 - テキストデコードを介さず座標をgroundingする設計。

4. どうやって有効だと検証した?

- SpatialLiDAR-QAデータセットで評価。 - 代表的なLiDAR–languageモデルおよびmulti-camera VLMと比較。 - 精密な座標予測タスクで大幅な改善を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている代表的なLiDAR–languageモデルおよびmulti-camera VLM。 - 関連手法としてLiDAR–languageモデル、multi-camera VLM、LLM for autonomous driving。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang

分類: cs.CV, cs.RO

原文アブストラクト

LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.

関連論文

PR本紙発行元 EmplifAI