日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.20586

CoRef-GS: マルチエージェント協調型参照3Dガウシアンスプラッティング

CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

シェア:XThreadsFacebookLINEはてブBluesky

複数ロボットが各自構築した3Dガウシアンマップを統合し、言語クエリで対象物を特定する協調型シーン理解フレームワークを提案。

詳しい要約

1. どんなもの?

- 複数エージェントの協調的な referring scene understanding を扱う CoRef-GS を提案。 - 各エージェントが独立に構築した open-vocabulary instance-aware Gaussian map を統合し、融合地図上で言語クエリを grounding する。 - 対象や landmark が他エージェントの観測に由来しても、問い合わせロボットの視点から空間関係を解釈する必要がある。 - この問題を cooperative referring Gaussian grounding over fused maps として定式化。 - 二足歩行ロボット2体の実世界・シミュレーション室内ベンチマーク CoQuad-Ref も導入。

2. 先行研究と比べてどこがすごい?

- 既存の language-aware Gaussian 手法は主に単一地図のクエリを対象とし、協調設定を扱わない。 - Gaussian registration 手法は幾何・測光整合を最適化するが、言語 grounding 指向の意味的互換性を保たない。 - CoRef-GS は幾何整合性・インスタンス意味比較性・視点条件付き関係推論を同時に要求する設定に対応。 - 実世界 referring mIoU を ReferSplat の 52.6% から 68.8% へ改善。 - シミュレーションで回転誤差を粗初期化 2.58° から refinement 後 0.15° へ低減。

3. 技術・手法の肝は?

- 各エージェントで open-vocabulary instance-aware Gaussian map を構築。 - 部分重複地図を cross-agent alignment module で幾何・意味一貫性に基づき整列。 - クエリは view-conditioned mask relation graph を用いて grounding。 - 融合地図上で対象・文脈 landmark が他エージェント由来でも、問い合わせ視点の空間関係を解釈。 - 詳細な損失設計や学習手順は要旨からは不明。

4. どうやって有効だと検証した?

- 二足歩行ロボット2体の実世界・シミュレーション室内シーンからなる CoQuad-Ref ベンチマークを構築。 - シミュレーションで回転誤差が粗初期化 2.58° から refinement 後 0.15° に低減することを確認。 - 実世界で referring mIoU が ReferSplat の 52.6% から 68.8% へ向上することを確認。 - ベンチマークとソースコードは公開予定。

5. 議論はある?

- 協調設定では対象や landmark が他エージェント観測に由来し、問い合わせ視点からの関係解釈が必要という課題を指摘。 - 既存手法は単一地図クエリか幾何・測光整合に偏り、言語 grounding の意味互換性を欠くと議論。 - 限界や失敗事例、計算コスト、スケーラビリティに関する議論は要旨からは不明。

6. 次に読むべき論文は?

- ReferSplat(比較対象の referring Gaussian 手法)。 - language-aware Gaussian 手法(単一地図クエリ系)。 - Gaussian registration 手法(幾何・測光整合系)。 - open-vocabulary instance-aware Gaussian map 構築に関する研究。 - 協調 SLAM・multi-agent 3D scene understanding の関連研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhikun Zhou, Kunyu Peng, Runyi Yang, Junhao Cai, Di Wen, Ruiping Liu, Danda Pani Paudel, Yi Zhou, Luc Van Gool, Kailun Yang

分類: cs.RO, cs.CV, eess.IV

原文アブストラクト

Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independently reconstructed maps are aligned and fused. In this setting, the referred target or its contextual landmark may come from another agent's observations, while spatial relations must still be interpreted from the querying robot's viewpoint. We formulate this problem as cooperative referring Gaussian grounding over fused maps, which requires geometric alignability, instance-level semantic comparability, and view-conditioned relation reasoning. Existing language-aware Gaussian methods mainly focus on single-map querying, whereas Gaussian registration methods optimize geometric or photometric alignment without preserving language-grounding-oriented semantic compatibility. We propose CoRef-GS, a cooperative referring Gaussian splatting framework. CoRef-GS constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph. We further introduce CoQuad-Ref, a dual-quadruped benchmark spanning both real-world and simulated indoor scenes. Experiments show that, on simulated scenes, CoRef-GS reduces the rotation error from 2.58° after coarse initialization to 0.15° after refinement, and improves real-world referring mIoU over ReferSplat from 52.6% to 68.8%. The established benchmark and source code will be publicly released at https://github.com/ruojiruoli17/CoRef-GS.git.

関連論文

PR本紙発行元 EmplifAI