日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
深度推定arXiv:2606.15681

自己教師あり単眼ビデオ深度推定のための3D一貫性最適化

3D Consistency Optimization for Self-Supervised Monocular Video Depth Estimation

シェア:XThreadsFacebookLINEはてブBluesky

ビデオフレームを独立に扱う従来法の問題を解決するため、深度推定を多視点3D再構成問題として捉え、3D一貫性最適化フレームワークを導入した。画像レベルのフォトメトリックレンダリング、ワールド座標の幾何学的整合、多スケール時間勾配一貫性の3つの制約でフレームを統一的に最適化し、自己教師あり学習とゼロショット臨床環境で最先端の精度を達成した。

著者: Yuanye Liu, Ke Zhang, Junzhe Jiang, Li Zhang, Vishal Patel, Xiahai Zhuang

分類: cs.CV

原文アブストラクト

Reliable monocular video depth estimation is crucial for downstream 3D reasoning and embodied AI in endoscopic navigation. However, existing self-supervised approaches typically treat video frames independently or rely on weak temporal regularization. These methods, lacking a holistic perception of the underlying 3D scene, inevitably suffer from geometrically inconsistent predictions and severe cross-frame drift. To address these limitations, we introduce a new paradigm that recasts sequential video depth estimation as an unconstrained multi-view 3D reconstruction problem, enabling full exploitation of the powerful geometric priors embedded in recent 3D foundation models. The core of our approach is a 3D consistency optimization framework driven by three constraints: image-level photometric rendering, explicit world-coordinate geometric alignment, and multi-scale temporal gradient consistency. Such unified optimization elegantly anchors isolated frames to a globally coherent 3D structure. Our method has been validated in both the self-supervised training scenarios and challenging zero-shot clinical environments. Results show that the proposed approach achieves state-of-the-art spatial accuracy, outperforming the frame-based, video-based depth estimators and the multi-view 3D reconstruction baselines.

関連論文