日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3DマッピングarXiv:2610.09518

ActiveLang: 意味的不確かさに導かれる探索による能動的オープンボキャブラリ3Dマッピング

ActiveLang: Active Open-Vocabulary 3D Mapping with Semantic-Uncertainty-Guided Exploration

シェア:XThreadsFacebookLINEはてブBluesky

意味的不確かさを手がかりに視点を選び、言語注釈付き3Dマップを効率的に構築する能動的オープンボキャブラリマッピングシステムを提案。

詳しい要約

1. どんなもの?

- ロボットが未知環境で多様なタスクを行うために、言語注釈付き3Dマップを自律的に構築するシステム。 - オープンボキャブラリなシーン理解と人間-ロボットインタラクションをサポート。 - 意味的不確実性に基づく探索を行い、オンラインで言語特徴を適応。 - コンパクトな dual-Gaussian 表現を用いて、形状・外観・オープンボキャブラリ意味を再構成。

2. 先行研究と比べてどこがすごい?

- オンラインおよびオフラインのベースラインと比較して、2Dおよび3Dオープンボキャブラリセグメンテーションで大幅な改善。 - より少ない観測と低い計算コストで効果的なマッピングを実現。 - 能動的探索により、言語注釈付き3Dマップをより効率的に構築できることを示唆。

3. 技術・手法の肝は?

- コンパクトな dual-Gaussian 表現上でオンライン言語特徴適応を実行。 - 意味的不確実性をガイドとして、情報量の多い視点を効率的に選択するプランナー。 - 形状・外観・オープンボキャブラリ意味を共同で再構成し、メモリオーバーヘッドを抑える。

4. どうやって有効だと検証した?

- Replica および ScanNet++ データセットで実験。 - 2Dおよび3Dオープンボキャブラリセグメンテーションの性能を評価。 - オンラインおよびオフラインのベースラインと比較し、大幅な改善を確認。

5. 議論はある?

- 能動的探索が言語注釈付き3Dマップ構築の効率を高めることを強調。 - 具体的な議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- Replica および ScanNet++ データセットを用いた研究。 - オンラインおよびオフラインのオープンボキャブラリ3Dマッピングのベースライン手法。 - dual-Gaussian 表現や言語特徴適応に関する関連研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Liyan Chen, Hairong Yin, Huangying Zhan, Yi Xu, Raymond A. Yeh, Philippos Mordohai

分類: cs.CV, cs.RO

原文アブストラクト

As robots increasingly assist humans with diverse tasks, they need both geometric and semantic understanding of their surroundings. Moreover, robots often operate in unfamiliar environments and take on new tasks without knowing the relevant concepts ahead of time. This motivates language-annotated 3D maps that support open-vocabulary scene understanding and human-robot interaction. We introduce ActiveLang, an autonomous system for active open-vocabulary 3D mapping with semantic-uncertainty-guided exploration. ActiveLang performs online language-feature adaptation on a compact dual-Gaussian representation to jointly reconstruct scene geometry, appearance, and open-vocabulary semantics with modest memory overhead. Its planner efficiently selects informative viewpoints, enabling effective mapping with fewer observations and lower computational cost. Experiments on Replica and ScanNet++ demonstrate substantial improvements in 2D and 3D open-vocabulary segmentation over both online and offline baselines, highlighting that actively exploring scenes builds language-annotated 3D maps more efficiently.

関連論文

PR本紙発行元 EmplifAI