日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D再構成/医療ロボティクスarXiv:2609.23961

Colon3R: 単眼内視鏡動画からのクロスドメイン3D再構成

Colon3R: Cross-Domain 3D Reconstruction from Monocular Colonoscopic Video

シェア:XThreadsFacebookLINEはてブBluesky

ファントムやシミュレーションデータで学習した幾何基盤モデルを、ラベルなしの生体内大腸内視鏡動画に半教師あり適応させ、深度・点群・カメラ姿勢を高精度に推定するフレームワークを提案。

詳しい要約

1. どんなもの?

- 単眼内視鏡動画からの3D再構成を目的としたフレームワーク - Colon3Rと命名 - クロスドメイン半教師あり学習 - 事前学習済みVGGTを基盤 - ラベル付きファントム・シミュレーションデータからラベルなしin-vivo大腸内視鏡へ転移 - カメラ・深度・ポイントマップの幾何を同時に推定

2. 先行研究と比べてどこがすごい?

- 従来の多視点3D再構成は安定した対応と剛体近似に依存 - 内視鏡特有の弱テクスチャ・鏡面反射・視野重複不足・非剛体運動で破綻 - 既存内視鏡手法はドメイン固有の教師あり学習に依存 - in-vivoラベル不足で幾何基盤モデルの適応が困難 - Colon3Rはターゲットドメインの幾何アノテーション不要 - ソースのみのファインチューニングと異なり、ラベルなしin-vivo動画を教師由来のクロスビュー監督で直接活用

3. 技術・手法の肝は?

- 事前学習済みVGGTを基盤としたクロスドメイン半教師ありフレームワーク - 階層的準剛体信頼性(hierarchical quasi-rigid reliability)を提案 - シーケンス・有向ペア・ピクセルの3レベルで信頼できる監督を選択 - ソース保持適応(source-preserving adaptation) - ターゲットドメイン適応中に学習済み結合幾何を保持 - 教師由来のクロスビュー監督を利用

4. どうやって有効だと検証した?

- 深度・ポイントマップ・カメラポーズ推定で最先端手法を上回る総合性能を実証 - 実in-vivo大腸内視鏡の定性的比較で、臨床ドメインシフト下でもより完全で幾何的に一貫した再構成を確認 - 詳細な実験設定は要旨からは不明

5. 議論はある?

- 臨床ドメインシフト下での再構成の完全性と幾何的一貫性が向上 - コードは論文採択後に公開予定 - 限界や失敗事例、計算コスト、一般化可能性に関する議論は要旨からは不明

6. 次に読むべき論文は?

- VGGT(Visual Geometry Grounded Transformer) - 従来の多視点3D再構成手法 - 既存の内視鏡3D再構成手法 - 幾何基盤モデル(geometry foundation models) - 半教師ありドメイン適応 - 準剛体シーン再構成

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhihao Xing, Yingyu Wang, Liang Zhao, Shoudong Huang

分類: cs.CV, cs.RO

原文アブストラクト

Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue motion. Conventional multi-view 3D reconstruction methods rely on stable correspondences and approximate rigidity, which are often violated in colonoscopy. Existing endoscopic methods often rely on domain-specific supervision, whereas there are not enough in-vivo labeled data available to adapt geometry foundation models to clinical colonoscopy. We present Colon3R, a cross-domain semi-supervised framework built on pretrained VGGT that transfers coupled camera, depth, and pointmap geometry from labeled phantom and simulated data to unlabeled in-vivo colonoscopy without requiring target-domain geometric annotations. Unlike source-only fine-tuning, which learns only from phantom and simulated data, Colon3R directly exploits unlabeled in-vivo video through teacher-derived cross-view supervision. Our proposed hierarchical quasi-rigid reliability selects reliable supervision at the sequence, directed-pair, and pixel levels, while source-preserving adaptation retains the learned coupled geometry during target-domain adaptation. Extensive experiments demonstrate that our method achieves superior overall performance over state-of-the-art approaches in depth, pointmap, and camera pose estimation. Qualitative comparisons on real in-vivo colonoscopy further show substantially more complete and geometrically consistent reconstructions than competing methods under clinical domain shift. The code will be public available after the paper is accepted.

PR本紙発行元 EmplifAI