日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3Dアバター編集arXiv:2610.03599

ManifoldSplat: 言語ガイドによる3Dガウシアンヘッドアバターの意味的形状編集

ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars

シェア:XThreadsFacebookLINEはてブBluesky

単眼動画から再構成したアニメーション可能な3Dガウシアンヘッドアバターに対し、FLAMEマニフォールド上で言語指示による局所的な形状編集を高速に行うフレームワークを提案。

詳しい要約

1. どんなもの?

- 単眼動画から再構成した animatable 3D Gaussian Splatting ヘッドアバターに対し、自然言語で局所的な形状編集を行う初の end-to-end フレームワーク。 - 編集を非構造な Gaussian cloud 上で直接行わず、構造化された FLAME manifold 内で実行する。 - これにより identity と animation を厳密に保持する。 - 約90秒で再構成・編集し、consumer GPU で約800 FPS 描画。

2. 先行研究と比べてどこがすごい?

- 従来の text-driven 手法は fine-grained な局所制御が難しく、特徴が絡み合い幾何的一貫性を欠く。 - 自然言語による形状変更は per-prompt 最適化が遅く、identity や rigging を損なうことが多い。 - 本手法は FLAME manifold 内で編集し、identity と animation を厳密に保持。 - 局所的な prompt alignment、幾何的一貫性、identity 保持で新たな state-of-the-art を達成。

3. 技術・手法の肝は?

- 編集を構造化された FLAME manifold 内で行い、非構造な Gaussian cloud を直接最適化しない。 - DeltaRegion を導入:領域ごとに disentangle された Conditional Variational Autoencoder (CVAE) が feedforward で shape delta を出力。 - さらに refining stage により view-consistent な詳細を復元。 - これらを組み合わせ、単眼動画から再構成した animatable 3D Gaussian Splatting アバターを編集。

4. どうやって有効だと検証した?

- Extensive evaluations を実施し、局所的な prompt alignment、幾何的一貫性、identity 保持で新たな state-of-the-art を示した。 - 約90秒で再構成・編集、consumer GPU で約800 FPS 描画という性能も報告。 - 具体的なデータセット名や評価指標は要旨からは不明。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗ケース、議論の詳細には言及されていない。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない。 - 関連手法として 3D Gaussian Splatting、FLAME、Conditional Variational Autoencoder (CVAE)、text-driven manipulation が挙げられる。 - 同分野の定番として 3D Gaussian Splatting や FLAME ベースの head avatar 研究を次に読むべき。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Antonio Canela, Jordi Sànchez-Riera

分類: cs.CV

原文アブストラクト

High-fidelity 3D head avatars have reached near-photorealistic quality. While recent methods enable text-driven manipulation, they struggle to provide fine-grained localized control, often entangling features or lacking geometric consistency. Modifying geometry through natural language currently requires slow per-prompt optimization or compromises identity and rigging. We present ManifoldSplat, the first end-toend framework for language-guided semantic shape editing of animatable 3D Gaussian Splatting avatars reconstructed from monocular videos. By performing edits within the structured FLAME manifold rather than directly optimizing an unstructured Gaussian cloud, we strictly preserve identity and animation. We introduce DeltaRegion, a per-region disentangled Conditional Variational Autoencoder (CVAE) delivering feedforward shape deltas, alongside a refining stage to recover view-consistent details. ManifoldSplat reconstructs and edits an avatar in ~90 seconds on a consumer GPU, rendering at ~800 FPS. Extensive evaluations demonstrate our approach sets a new state-of-the-art in localized prompt alignment, geometric coherence, and identity preservation. Project page and code: https://a-canela.github.io/manifoldsplat/

PR本紙発行元 EmplifAI