日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D検出arXiv:2610.03015

OmniAct3D: 全天球3D検出のための基盤幾何と根拠付き推論の活用

OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

シェア:XThreadsFacebookLINEはてブBluesky

透視画像で学習済みの視覚基盤モデルを全天球(ERP)画像に適応させ、球面レイ幾何アダプタと視覚-行動推論チェーンにより360度の3D物体検出精度を向上させた研究。

詳しい要約

1. どんなもの?

- モバイルembodied agent向けのpanoramic 3D detectionフレームワーク - Vision Foundation Models (VFMs)の視覚・幾何priorをERPに適応 - ERP-Ray Geometry Adapter (ERGA-Ray)、Visual-Action Reasoning Chain (VARC)、Appearance-Guided Heading Expert (AGHE)を統合 - 単一のequirectangular projection (ERP)画像で360度の連続シーンを扱う

2. 先行研究と比べてどこがすごい?

- 既存のVFMベース3D検出器は狭視野monocular画像や離散的なperspective viewに依存 - それらはcoherent surround perceptionを制限 - OmniAct3DはERPに適応しつつVFM priorを保持 - Spheriverseで前最高3D検出器を2.96 NDSポイント上回る - PanoMMOccで未適応VFMベースラインを24.87 mAPポイント上回る

3. 技術・手法の肝は?

- ERGA-Ray: 球面viewing rayと周期的空間構造をモデル化し幾何ミスマッチを解消 - VARC: 各hypothesisをpanoramic evidenceにgroundingし構造化geometric actionに変換 - AGHE: 固定token budgetで失われる局所cueを高解像度再エンコードで回復しheading推定 - perspective-trained VFM検出器をERPに適応

4. どうやって有効だと検証した?

- Spheriverseで前最高3D検出器より2.96 NDSポイント改善 - PanoMMOccで未適応VFMベースラインより24.87 mAPポイント改善 - target-specific geometry adaptationによりVARCは同構成mAPの95–98%を保持 - 異なるsensing configuration間で再利用可能なobject-level 3D reasoningを示す

5. 議論はある?

- ERPは幾何と視覚情報の組織が異なり、object-relevant cueのモデル化・定位・保持が困難 - 直接転移は難しい - VARCのobject-level 3D reasoningがsensing configurationをまたいで再利用可能である点を議論 - その他の限界や議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 既存のVFM-based 3D detectors、perspective-trained VFM detectors、Spheriverse、PanoMMOcc - 関連手法: Vision Foundation Models (VFMs)、equirectangular projection (ERP) - 同分野の定番: monocular 3D detection、panoramic 3D detection

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection (ERP) instead encodes a continuous 360 scene in a single image. Direct transfer remains difficult because ERP organizes geometry and visual information differently, making object-relevant cues hard to model, localize, and preserve. We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. To resolve geometric mismatch, the ERP-Ray Geometry Adapter (ERGA-Ray) models spherical viewing rays and periodic spatial structure. To localize evidence in scene-wide context, the Visual-Action Reasoning Chain (VARC) grounds each hypothesis in relevant panoramic evidence and converts it into a structured geometric action. To recover local cues lost under fixed token budgets, the Appearance-Guided Heading Expert (AGHE) re-encodes object regions at higher resolution for heading estimation. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. With target-specific geometry adaptation, VARC retains 95--98% of the same-configuration mAP, indicating reusable object-level 3D reasoning across sensing configurations. The source code will be made publicly available at https://github.com/FeiT-FeiTeng/OmniAct3D.

関連論文

PR本紙発行元 EmplifAI