日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
物体ナビゲーションarXiv:2609.27218

NaviScale: 物体ナビゲーション向け大規模セマンティックマップデータセットの生成

NaviScale: Generating Large-Scale Semantic Map Datasets for Object Navigation

シェア:XThreadsFacebookLINEはてブBluesky

実環境の3D再構成なしに、実住宅の間取り図と部屋単位のセマンティック・障害物マップを組み合わせて大規模なセマンティックマップ訓練データを生成し、物体ナビゲーションの性能を向上させる手法を提案。

詳しい要約

1. どんなもの?

- 意味マップベースの ObjectNav 用の大規模データ生成フレームワーク NaviScale を提案。 - 予測器は partial と complete の semantic map のペアで学習でき、各サンプルごとに完全な 3D 環境を再構成する必要がない。 - 実住宅の floorplan に MP3D と HM3DSem から抽出した room-level の semantic/obstacle map を組み合わせて生成。 - 192,000 の semantic map を 24,000 floorplan・12,794 property から生成。

2. 先行研究と比べてどこがすごい?

- 実 3D 環境からの大量アノテーション収集は困難という課題に対し、完全 3D 再構成を不要にした点が新しい。 - 既存の ObjectNav データセットと比べ、floorplan と room map の組合せで多様性を拡大。 - 予測アーキテクチャを変えずに HM3D で 64.3% SR / 34.8% SPL、MP3D で 43.1% SR / 16.8% SPL を達成。 - 具体的な先行研究名は要旨からは不明。

3. 技術・手法の肝は?

- floorplan と MP3D/HM3DSem 由来の room-level semantic/obstacle map を合成。 - inter-room scaling で floorplan レベルの構造的多様性を増加。 - intra-room scaling で同一 floorplan にカテゴリ一致した異なる room map を充填。 - Visibility through Ray Casting (VisRC) で FOV・sensing range・occlusion を考慮した partial observation に変換。

4. どうやって有効だと検証した?

- 300k 学習イテレーションで HM3D と MP3D の ObjectNav ベンチマークを評価。 - 合成マップの品質、semantic-segmentation 誤差の影響、実機ロボットへの展開を追加実験で検証。 - 予測アーキテクチャを変更せずに上記性能を確認。

5. 議論はある?

- 合成マップの品質評価と semantic-segmentation 誤差の影響を議論。 - 実機ロボットへの展開可能性を検討。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- MP3D (Matterport3D) と HM3DSem を参照。 - ObjectNav の関連研究として、semantic-map-based ObjectNav や HM3D/MP3D ベンチマークの定番手法を次に読むべき。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chuanlin Lan, Yanwei Zheng, Weijian Liu, Zhitong Zhou, Xiao Zhang, Fuzhen Zhuang, Dongxiao Yu

分類: cs.RO

原文アブストラクト

Embodied navigation requires spatial representations that generalize across unseen environments, yet collecting large amounts of annotated data from real 3D environments is difficult. We propose NaviScale for semantic-map-based object navigation (ObjectNav), whose predictor can be trained on pairs of partial and complete semantic maps without reconstructing a complete 3D environment for every training sample. The framework generates large-scale semantic map training data by composing floorplans of real homes with room-level semantic and obstacle maps extracted from MP3D and HM3DSem. NaviScale increases data diversity in two ways: inter-room scaling increases floorplan-level structural diversity, while intra-room scaling fills each fixed floorplan with different combinations of room maps matched by room category. Visibility through Ray Casting (VisRC) converts the composed maps into partial observations that account for field of view, sensing range, and occlusion. The resulting dataset contains 192,000 semantic maps generated from 24,000 floorplans associated with 12,794 properties. With 300k training iterations and the training and inference settings described in this paper, the system reaches 64.3% SR and 34.8% SPL on HM3D, together with 43.1% SR and 16.8% SPL on MP3D, without changing the prediction architecture. Additional experiments evaluate the quality of the composed maps, the effects of semantic-segmentation errors, and deployment on a physical robot.

関連論文

PR本紙発行元 EmplifAI