日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
シーングラフarXiv:2610.07569

OpenSplatGraph: 密なセマンティックマップから構造化シーングラフへ

OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception

シェア:XThreadsFacebookLINEはてブBluesky

3Dガウシアンスプラッティングによるオープンボキャブラリなセマンティックマップから、物体と関係性を永続的に保持する3Dシーングラフを構築するフレームワークを提案。

詳しい要約

1. どんなもの?

- 3D Gaussian Splatting ベースのオンライン open-vocabulary semantic map から、永続的な 3D scene graph を直接構築する統合フレームワーク OpenSplatGraph を提案。 - 密な semantic map に reliability-aware semantic field を付加し、confidence-aware かつ query-conditioned な object extraction を可能にする。 - 抽出した object instance を永続的な graph node に紐付け、観測や query をまたいで属性と関係を逐次更新する。 - language-guided object grounding と structured relational reasoning の両方を、Gaussian ベースの幾何忠実度を保ちつつ支援する。

2. 先行研究と比べてどこがすごい?

- 従来の 3D Gaussian Splatting ベース mapping は高忠実な幾何と効率的な open-vocabulary perception を実現するが、semantics を非構造な feature field として表現するため object-centric reasoning が制限される。 - 一方、3D scene graph は object と関係を明示的にモデル化するが、密な semantic map を十分活用しない疎な幾何表現から構築されることが多い。 - 本研究は密な semantic mapping と永続的な object-centric 表現を密結合し、両者の利点を統合する点が新しい。

3. 技術・手法の肝は?

- オンラインの Gaussian ベース open-vocabulary semantic map を基盤とする。 - reliability-aware semantic field を導入し、軽量な observation statistics を保持して confidence-aware な query-conditioned object extraction を行う。 - 抽出された object instance を永続的な graph node に関連付け、観測・query をまたいで object 属性と関係を逐次更新する。 - 密な semantic mapping と永続的な object-centric 表現を密結合することで、language-guided object grounding と structured relational reasoning を両立する。

4. どうやって有効だと検証した?

- 標準的な 3D scene understanding benchmarks と実世界ロボティクス実験で包括的に評価。 - オンライン open-vocabulary perception および下流のロボティクスタスクにおいて競争力のある性能を示すと報告。 - 具体的なベンチマーク名や指標、実験設定の詳細は要旨からは不明。

5. 議論はある?

- 要旨からは不明。 - 想定される論点として、reliability-aware semantic field の confidence 推定精度、graph node の永続性・更新の一貫性、密な map と graph の結合に伴う計算コスト、実世界環境でのスケーラビリティなどが考えられるが、要旨には明示されていない。

6. 次に読むべき論文は?

- 3D Gaussian Splatting ベースの semantic mapping(例: Gaussian Splatting を用いた open-vocabulary 3D mapping 手法)。 - 3D scene graph 構築手法(例: 疎な幾何からの scene graph 生成)。 - open-vocabulary 3D perception および language-guided object grounding に関する研究。 - 具体的な参照論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Binh Long Nguyen, Kien Nguyen, Clinton Fookes, Peyman Moghadam

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

Dense 3D mapping with semantic understanding is essential for robotic perception in complex environments. Recent 3D Gaussian Splatting-based mapping approaches enable high-fidelity geometry and efficient open-vocabulary perception, but typically represent semantics as unstructured feature fields that limit object-centric reasoning. In contrast, 3D scene graphs explicitly model objects and their relationships for structured reasoning, but are commonly constructed from sparse geometric representations that do not fully exploit dense semantic maps. In this work, we present OpenSplatGraph, a unified framework that constructs persistent 3D scene graphs directly from an online Gaussian-based open-vocabulary semantic map. The proposed framework augments the dense semantic map with a reliability-aware semantic field that maintains lightweight observation statistics for confidence-aware, query-conditioned object extraction. Extracted object instances are associated with persistent graph nodes, allowing object attributes and relationships to be incrementally updated across observations and queries. By tightly coupling dense semantic mapping with persistent object-centric representations, our framework supports both language-guided object grounding and structured relational reasoning while preserving the geometric fidelity of Gaussian-based mapping. Comprehensive evaluations on standard 3D scene understanding benchmarks and real-world robotic experiments demonstrate that OpenSplatGraph achieves competitive performance for online open-vocabulary perception and downstream robotic tasks. Project page: https://csiro-robotics.github.io/OpenSplatGraph.

関連論文

PR本紙発行元 EmplifAI