生成モデルarXiv:2406.01056
仮想アバター生成モデルによる世界ナビゲーション
Virtual avatar generation models as world navigators
拡散トランスフォーマーで人間のクライミング動作を動画として生成するSABR-CLIMBを提案し、大規模データセットNAV-22Mで汎用仮想アバターの可能性を示した。
著者: Sai Mandava
分類: cs.CV, cs.AI, cs.HC, cs.LG, cs.RO
原文アブストラクト
We introduce SABR-CLIMB, a novel video model simulating human movement in rock climbing environments using a virtual avatar. Our diffusion transformer predicts the sample instead of noise in each diffusion step and ingests entire videos to output complete motion sequences. By leveraging a large proprietary dataset, NAV-22M, and substantial computational resources, we showcase a proof of concept for a system to train general-purpose virtual avatars for complex tasks in robotics, sports, and healthcare.