VGGTが知る重なり: 幾何基盤モデルにおける共可視性の探査
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
VGGTの内部表現が共可視性を暗黙に符号化していることを発見し、軽量なMixture-of-Expertsヘッドを追加してRGB画像のみから共可視性を高精度に分類する手法を提案した。
著者: Filippo Ziliotto, Luciano Serafini, Lamberto Ballan, Tommaso Campari
分類: cs.CV, cs.AI
原文アブストラクト
A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios with minimal overlap. We demonstrate that VGGT implicitly encodes co-visibility as an emergent behavior: without any supervision for this task, its internal representations exhibit a clear hierarchical structure mirroring that of large language models, i.e. early layers build a 3D-aware scene representation, while late layers act as dedicated co-visibility reasoners. In particular, we identify layer L17 as a negative anchor that consistently routes non-co-visible pairs for this backbone, regardless of the evaluation setting, providing task-grounded evidence of layer specialization in a geometry-grounded foundation model. Building on this, we introduce Co-VGGT, which freezes VGGT and trains only a lightweight layer-wise mixture-of-experts head (less than 7.5M parameters) to classify co-visibility from RGB alone, treating each layer as a specialized expert whose geometric abstraction is adaptively weighted per input pair. On the Co-VisiON benchmark, Co-VGGT surpasses the human annotation baseline and improves over prior work by more than 25% pairwise and 10% multiview. Pairwise predictions are well-calibrated (ECE=0.030), enabling direct use as edge weights in visibility graphs for downstream SfM and SLAM pipelines without post-hoc correction. Code and data are available.
関連論文
- 大規模再構成モデルを用いた人と物体のインタラクション再構成3D再構成
- PIVOT: 実世界3D再構成における姿勢・内部パラメータ・新視点評価のためのマルチ軌道データセットとテストベッド3D再構成
- OccamView: フレーム予算制約下のアクティブ3Dガウス再構成のためのオブジェクト条件付き視点選択3D再構成
- DerainSplat: スパースな雨天視点からのフィードフォワードによるクリーンな3Dガウススプラッティング3D再構成
- Stipple: 視覚慣性トラッキングによるリアルタイムインクリメンタルガウシアンスプラッティング3D再構成
- DA-NBV:船舶の効率的な3D再構成のための方向認識型次善視点プランナー3D再構成