RGB-D-Fusion: 人物深度マップを生成するマルチモーダル拡散モデル
RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
低解像度の単眼RGB画像から人物の高解像度深度マップを生成する2段階の拡散モデルを提案し、深度ノイズ拡張で超解像の頑健性を高めた。
著者: Sascha Kirch, Valeria Olyunina, Jan Ondřej, Rafael Pagés, Sergio Martin, Clara Pérez-Molina
分類: cs.CV, cs.LG
原文アブストラクト
We present RGB-D-Fusion, a multi-modal conditional denoising diffusion probabilistic model to generate high resolution depth maps from low-resolution monocular RGB images of humanoid subjects. RGB-D-Fusion first generates a low-resolution depth map using an image conditioned denoising diffusion probabilistic model and then upsamples the depth map using a second denoising diffusion probabilistic model conditioned on a low-resolution RGB-D image. We further introduce a novel augmentation technique, depth noise augmentation, to increase the robustness of our super-resolution model.