NaviScale: 物体ナビゲーション向け大規模セマンティックマップデータセットの生成
NaviScale: Generating Large-Scale Semantic Map Datasets for Object Navigation
実環境の3D再構成なしに、実住宅の間取り図と部屋単位のセマンティック・障害物マップを組み合わせて大規模なセマンティックマップ訓練データを生成し、物体ナビゲーションの性能を向上させる手法を提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Chuanlin Lan, Yanwei Zheng, Weijian Liu, Zhitong Zhou, Xiao Zhang, Fuzhen Zhuang, Dongxiao Yu
分類: cs.RO
原文アブストラクト
Embodied navigation requires spatial representations that generalize across unseen environments, yet collecting large amounts of annotated data from real 3D environments is difficult. We propose NaviScale for semantic-map-based object navigation (ObjectNav), whose predictor can be trained on pairs of partial and complete semantic maps without reconstructing a complete 3D environment for every training sample. The framework generates large-scale semantic map training data by composing floorplans of real homes with room-level semantic and obstacle maps extracted from MP3D and HM3DSem. NaviScale increases data diversity in two ways: inter-room scaling increases floorplan-level structural diversity, while intra-room scaling fills each fixed floorplan with different combinations of room maps matched by room category. Visibility through Ray Casting (VisRC) converts the composed maps into partial observations that account for field of view, sensing range, and occlusion. The resulting dataset contains 192,000 semantic maps generated from 24,000 floorplans associated with 12,794 properties. With 300k training iterations and the training and inference settings described in this paper, the system reaches 64.3% SR and 34.8% SPL on HM3D, together with 43.1% SR and 16.8% SPL on MP3D, without changing the prediction architecture. Additional experiments evaluate the quality of the composed maps, the effects of semantic-segmentation errors, and deployment on a physical robot.
関連論文
- Hydra-Nav: 適応的二重プロセス推論による物体ナビゲーション物体ナビゲーション
- ユーザー中心物体ナビゲーション:個人の習慣を統合したパーソナライズド物体探索ベンチマーク物体ナビゲーション
- 3DGSNav: アクティブ3Dガウシアンスプラッティングによる視覚言語モデルの物体ナビゲーション推論強化物体ナビゲーション
- PIGEON: VLM駆動の関心点選択による物体ナビゲーション物体ナビゲーション
- PersONAL: パーソナライズされた身体性エージェントのための包括的ベンチマーク物体ナビゲーション
- osmAG-LLM: セマンティックマップと大規模言語モデルの推論によるゼロショットオープンボキャブラリ物体ナビゲーション物体ナビゲーション