日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
協調知覚arXiv:2610.11577

CoCam4D: カメラのみの自動運転のための幾何学認識型協調4D知覚

CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving

シェア:XThreadsFacebookLINEはてブBluesky

複数車両が3Dガウス表現と不確かさを共有し、LiDARなしでカメラのみの協調知覚を実現するベイズフレームワークを提案。

詳しい要約

1. どんなもの?

- カメラのみの協調知覚フレームワーク - 複数車両が観測を共有し遮蔽や死角を補完 - Bayesian枠組みで幾何不確かさを明示的にモデル化 - VGGTベースの順伝播ネットワークで3D Gaussian表現と不確かさを生成 - コンパクトなGaussian primitivesを共有しLiDAR不要 - 通信向けに35バイトのDynamic Object Primitives (DOPs)を導入

2. 先行研究と比べてどこがすごい?

- 従来のカメラのみの協調知覚は単眼深度推定の不確かさに制約 - 提案手法は幾何不確かさをBayesianに扱い、他エージェントの信頼できる観測で深度不確かさを低減 - LiDARセンサを必要とせず、vision-only手法を一貫して上回る - OPV2V+で11.48%、DAIR-V2X-Cで10.62%の改善 - 具体的な先行研究名は要旨からは不明

3. 技術・手法の肝は?

- Bayesian協調知覚フレームワーク - VGGTベースの順伝播ネットワークで3D Gaussian scene representationsと不確かさ推定を生成 - 複数車両がコンパクトなGaussian primitivesを共有し観測を統合 - 他エージェントの信頼できる観測が深度不確かさを低減 - 実展開向けにDynamic Object Primitives (DOPs)を導入 - DOPsは35バイトのコンパクト表現でC-V2X通信に適合

4. どうやって有効だと検証した?

- 広範な実験を実施 - 最近のvision-only手法と比較し一貫して上回る - OPV2V+で11.48%の改善 - DAIR-V2X-Cで10.62%の改善 - 具体的な評価指標や実験設定の詳細は要旨からは不明

5. 議論はある?

- カメラのみの協調知覚における幾何不確かさの重要性を議論 - LiDAR不要の自動運転に向けた幾何に基づく協調知覚の可能性を示す - 限界や課題、今後の方向性についての具体的な議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法としてVGGTベースの3D Gaussian scene representation、Bayesian協調知覚、C-V2X通信、OPV2V+、DAIR-V2X-Cが挙げられる - 同分野の定番としてmulti-agent collaborative perception、camera-only 3D perception、monocular depth estimationの論文を読むべき

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Soham Pahari, Sudip Das, Arindam Das, Ujjwal Bhattacharya

分類: cs.RO, cs.CV

原文アブストラクト

Autonomous vehicles often suffer from limited perception due to occlusions, blind spots, limited sensor range, and the complex nature of surrounding environments. Multi-agent collaborative perception (CP) addresses these challenges by allowing vehicles to share sensory information and reconstruct the scene cooperatively. However, camera-only perception remains fundamentally limited by the uncertainty of distance-dependent monocular depth estimation. We propose CoCam4D, a Bayesian framework for collaborative perception that explicitly models geometric uncertainty. It uses a VGGT-based feedforward network to generate 3D Gaussian scene representations with associated uncertainty estimates, enabling multiple vehicles or agents to efficiently combine their observations. By sharing compact Gaussian primitives, reliable observations from one agent can reduce the depth uncertainty of another without requiring LiDAR sensors. To support real-world deployment, we introduce Dynamic Object Primitives (DOPs), a compact 35-byte representation designed for efficient C-V2X communication. Extensive experiments show that our proposed method consistently outperforms recent vision-only methods, achieving improvements of 11.48% on OPV2V+ and 10.62% on DAIR-V2X-C, demonstrating the potential of geometrically grounded collaborative perception for LiDAR-free autonomous driving.

関連論文

PR本紙発行元 EmplifAI