単一視点メッシュ再構成はロボットカメラの回転に汎化できるか?
Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?
ロボット搭載カメラの回転が単一視点メッシュ再構成の性能に与える影響を評価し、回転による深度推定やレイアウトの誤差を明らかにした。重力情報を活用した改良によりレイアウト誤差を大幅に削減した。
著者: Yu Zhan, Guangcheng Chen, Hanjing Ye, Zhiqin Cheng, Zanjia Tong, Wenjun Xu, Hong Zhang
分類: cs.CV, cs.RO
原文アブストラクト
Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reasoning and real-to-sim digital twins. However, robot-mounted cameras naturally rotate during manipulation and navigation, while learned single-view reconstruction models often rely on view-dependent priors and may generalize poorly to out-of-distribution camera rotations. Such rotations can introduce 3D inconsistencies, incorrect layouts, and violations of physical constraints, but this failure mode remains under-evaluated. We introduce an evaluation protocol with controlled axis-wise roll, pitch, and yaw sweeps to trace errors in monocular depth estimation (MDE), canonical object meshes, camera-space layout, and physical plausibility within a representative SAM3D-style pipeline. On the Aria Digital Twin dataset and a real Franka wrist-camera sequence, camera rotations induce MDE distortion, layout drift, and collision penetration, while canonical mesh predictions remain relatively stable. A two-stage SAM3D+FoundationPose pipeline is more robust than one-stage feed-forward layout prediction, and our Gravity-Aware Refinement reduces one-stage pairwise ICP-based layout-orientation error by 47.1$\%$. Our evaluation reveals that current single-view mesh reconstruction methods generalize poorly to robot camera rotation, and suggests that explicit gravity cues are important for reliable robotic single-view mesh reconstruction.
関連論文
- MV-dVRK: 空間的外科知覚のための多視点ベンチマーク3D再構成
- 大規模再構成モデルを用いた人と物体のインタラクション再構成3D再構成
- PIVOT: 実世界3D再構成における姿勢・内部パラメータ・新視点評価のためのマルチ軌道データセットとテストベッド3D再構成
- OccamView: フレーム予算制約下のアクティブ3Dガウス再構成のためのオブジェクト条件付き視点選択3D再構成
- DerainSplat: スパースな雨天視点からのフィードフォワードによるクリーンな3Dガウススプラッティング3D再構成
- Stipple: 視覚慣性トラッキングによるリアルタイムインクリメンタルガウシアンスプラッティング3D再構成