日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
シーン補完arXiv:2610.03031

CrowdOcc: 混雑屋内環境における四足歩行ロボットのための単眼意味論的シーン補完

CrowdOcc: Monocular Semantic Scene Completion for Quadruped Robots in Crowded Indoor Environments

シェア:XThreadsFacebookLINEはてブBluesky

混雑した屋内環境で四足歩行ロボット向けに、RGB-Dデータセットと単眼意味論的シーン補完フレームワークを提案し、遮蔽に頑健な幾何融合と人間中心の相互作用モデリングで最先端性能を達成した。

詳しい要約

1. どんなもの?

- 四足歩行ロボット向けの単眼Semantic Scene Completion (SSC) を、混雑した屋内環境で実現する研究。 - RGB-DデータセットCrowdOccと、単眼SSCフレームワークを提案。 - データセットは11の屋内シーンから25.1Kフレームを含み、static dynamic decouplingにより意味的占有アノテーションを構築。 - フレームワークはNormal Guided Scene Geometry Fusion (NGSGF)とHuman-Centric Sparse Interaction (HCSI)を組み合わせる。

2. 先行研究と比べてどこがすごい?

- 四足歩行ロボットの混雑屋内環境における単眼SSCは未開拓。 - 人間-シーンのオクルージョンが静的幾何を乱し、人間占有予測が不完全または空間的に誤配置される問題に対処。 - 提案手法はCrowdOccのscene-disjointテストセットでstate-of-the-artのSSC性能を達成。 - 15.80 IoU、11.40 mIoU、46.23 Human IoUを記録し、未見の屋内シーンへの汎化を示す。

3. 技術・手法の肝は?

- Normal Guided Scene Geometry Fusion (NGSGF): 深度認識リフティングを表面法線キューで補完し、オクルージョンに頑健な幾何を実現。 - Human-Centric Sparse Interaction (HCSI): 3Dで人間-人間および局所的な人間-シーン関係を選択的にモデル化。 - データセット構築: static dynamic decouplingにより意味的占有アノテーションを生成。

4. どうやって有効だと検証した?

- CrowdOccのscene-disjointテストセットで評価。 - 指標: 15.80 IoU、11.40 mIoU、46.23 Human IoUを達成。 - 未見の屋内シーンへの汎化を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の関連手法として、Monocular Semantic Scene Completion (SSC)やRGB-Dベースの3Dシーン理解、四足歩行ロボットのナビゲーションに関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Feiyang Chen, Jincheng Hu, Yiduo Chen, Jihao Li, Yue Liang, Bingzhao Gao, Yanjun Huang, Yuanjian Zhang

分類: cs.CV, cs.RO

原文アブストラクト

Monocular semantic scene completion (SSC) for quadruped robots remains underexplored in real crowded indoor environments, where human-scene occlusion disrupts static geometry and human occupancy predictions are often incomplete or spatially misplaced. We present CrowdOcc, an RGB-D dataset and monocular SSC framework for this setting. CrowdOcc contains 25.1K frames from 11 indoor scenes, with semantic occupancy annotations constructed through static dynamic decoupling. Our framework combines: (i) Normal Guided Scene Geometry Fusion (NGSGF) to complement depth-aware lifting with surface-normal cues for occlusion robust geometry; and (ii) Human-Centric Sparse Interaction (HCSI) to selectively model human-human and local human scene relations in 3D. Our method achieves state-of-the-art SSC performance on CrowdOcc's scene-disjoint test set, reaching 15.80 IoU, 11.40 mIoU, and 46.23 Human IoU, demonstrating generalization to unseen indoor scenes.

関連論文

PR本紙発行元 EmplifAI