日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.02803

LOCUS: 空間グラフを用いたランドマーク指向の容器識別

LOCUS: Landmark-Oriented Container Discrimination Using Spatial Graphs

シェア:XThreadsFacebookLINEはてブBluesky

シーングラフ上でCLIP埋め込みをGNNで更新し、物体の空間的・意味的関係を統合して目的物体が入っていそうな容器を推論する手法を提案。シミュレーションと実機で有効性を示した。

詳しい要約

1. どんなもの?

- ロボットが未構造環境で物体の空間的・意味的関係を推論するための手法 - 特に、目標物体が直接見えない場合に、どの容器を探索すべきかを判断する - Landmark-Oriented Container Discrimination Using Spatial Graphs (LOCUS) を提案 - 空間情報と意味情報を統合的に推論する - シミュレーションと実機で評価

2. 先行研究と比べてどこがすごい?

- 従来のCLIPによるコサイン類似度では、空間的・意味的に類似した容器を区別するのが困難 - ランダム、CLIP、Tidybot、LLMプランナーと比較して、多くの部屋クラスと構成で優位 - シーングラフのノード配置やラベルノイズに対して頑健性を示す - 実機モバイルマニピュレータで探索パイプラインを完遂

3. 技術・手法の肝は?

- 環境全体のシーングラフ上でGNNを訓練し、近接性に基づいてノード間で埋め込み情報を伝播 - CLIP埋め込みを更新し、空間情報と意味情報を統合 - 家庭用オントロジーの融合による意味知識でCLIP埋め込みを拡張 - シミュレータのメタデータからノイズのないシーングラフを生成し、検出に依存しない評価を実現

4. どうやって有効だと検証した?

- シミュレーションでノイズのないシーングラフを用いて評価 - ランダム、CLIP、Tidybot、LLMプランナーと比較し、多くの部屋クラスと構成で優位 - 実機モバイルマニピュレータで探索パイプラインを最初から最後まで実行 - Deticによるシーングラフラベルを含む実環境で頑健性を検証

5. 議論はある?

- シーングラフのノード配置やラベルノイズに対する頑健性を実機で確認 - 検出に依存しない評価手法を採用 - 具体的な限界や議論の詳細は要旨からは不明

6. 次に読むべき論文は?

- CLIP (Radford et al.) - Tidybot - LLM planner - Detic - シーングラフやGNNを用いたロボティクス研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Taylor Bergeron, Shibani Senthilbabu, Kevin Leahy

分類: cs.RO

原文アブストラクト

As robots are increasingly deployed in unstructured, real-world environments, the ability to reason about complex spatial and semantic relationships among objects remains a fundamental challenge in enabling robust and generalizable manipulation and navigation. For example, deciding where to search for an object that is not in plain sight depends on where it is physically plausible as well as semantically likely. Popular methods such as cosine similarity with CLIP struggle to disambiguate between spatially and semantically similar objects that could contain a target object. We propose Landmark-Oriented Container Discrimination Using Spatial Graphs (LOCUS). To jointly reason about spatial information and semantics, we train a GNN to update CLIP embeddings on a full environment scene graph by passing embedding information between nodes based on proximity. CLIP embeddings, augmented by semantic knowledge from a fusion of household ontologies provide a robust semantic signal. We evaluate in simulation on a noiseless scene graph from simulator metadata, making the approach detection-agnostic. Our approach outperforms random, CLIP, Tidybot, and an LLM planner in the majority of room classes and configurations in simulation. To show our approach is robust to scene graph node placement and label noise, we demonstrate a physical mobile manipulator running in an exploration pipeline start-to-finish including scene graph labels from Detic.

関連論文

PR本紙発行元 EmplifAI