日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D占有予測arXiv:2610.04356

SelectOccFlow: 3D占有とシーンフローのための選択的時空間集約

SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction

シェア:XThreadsFacebookLINEはてブBluesky

自動運転向けに、画像・時間・ボクセルの各領域で信頼できる情報だけを選択的に集約し、3D占有とシーンフローを高精度かつ頑健に予測するフレームワークを提案した。

詳しい要約

1. どんなもの?

- カメラベースの3D occupancyとscene flow予測のためのSelectOccFlowを提案。 - 画像・時間・voxel領域で選択的時空間集約を行うフレームワーク。 - 自動運転の包括的3Dシーン理解を目的とする。

2. 先行研究と比べてどこがすごい?

- 従来の集約は意味的に不整合な画像特徴、位置ずれした履歴観測、不完全なvoxel構造に敏感。 - SelectOccFlowはこれらを段階的に改善し、OpenOccでOccScore 44.9を達成し、従来最高を+4.2%改善。 - nuScenes-Cのcorruption下で平均OccScoreを+11.1%改善し、ロバスト性が向上。

3. 技術・手法の肝は?

- Semantic-Guided Sampling (SGS)で意味的priorに基づき特徴サンプリングを調整。 - State-Conditioned Temporal Aggregation (SCTA)でvoxel状態に応じて履歴証拠を選択的に取得。 - Extent-Aware Spatial Aggregation (ESA)で方向性構造サポートを活用し前景形状を洗練。

4. どうやって有効だと検証した?

- OpenOccでOccScore 44.9を達成し、従来最高を+4.2%改善。 - Occ3D-nusで競争力のあるoccupancy性能を維持。 - nuScenes-Cのcorruption下で平均OccScoreを+11.1%改善し、視覚的corruptionへのロバスト性を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- OpenOcc, Occ3D-nus, nuScenes-Cに関する研究。 - カメラベースのoccupancy予測とscene flow予測の関連手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuhang Wang, Kai Luo, Yuanfan Zheng, Kailun Yang

分類: cs.CV, cs.RO

原文アブストラクト

Comprehensive 3D scene understanding for autonomous driving requires modeling geometry, semantics, and motion. However, camera-based occupancy and scene flow prediction are sensitive to unreliable spatial and temporal aggregation, caused by semantically incompatible image features, misaligned historical observations, and incomplete voxel structures. To address this issue, we propose SelectOccFlow, a selective spatiotemporal aggregation framework that progressively refines contextual evidence across image, temporal, and voxel domains. To obtain semantically compatible image evidence, we design Semantic-Guided Sampling (SGS) to regulate feature sampling with semantic priors. Since reliable image evidence alone cannot resolve temporal inconsistency, we then present State-Conditioned Temporal Aggregation (SCTA) to selectively retrieve historical evidence according to voxel states. To further enhance the structural completeness of voxel representations, we introduce Extent-Aware Spatial Aggregation (ESA), which exploits directional structural support to refine foreground geometry. Experiments on OpenOcc demonstrate that SelectOccFlow achieves a state-of-the-art OccScore of 44.9, improving the previous best by +4.2%. It also maintains competitive occupancy performance on Occ3D-nus and improves the mean OccScore under nuScenes-C corruptions by +11.1%, demonstrating improved robustness to visual corruptions. The source code will be made publicly available at https://github.com/muchen1021/SelectOccFlow.

関連論文

PR本紙発行元 EmplifAI