UnfoldArt: テキストまたは画像からの完全な関節3Dオブジェクトのゼロショット復元
UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image
テキストや画像から、関節構造と隠れた内部形状を含む完全な3Dオブジェクトを復元する、エージェントベースの新しい手法を提案。
著者: Mohamed el Amine Boudjoghra, Ivan Laptev, Angela Dai
分類: cs.CV
原文アブストラクト
Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and motion from sparse observations remains challenging. Existing approaches remain largely constrained by lack of supervised data or lack the priors needed to reliably recover articulation, hidden geometry, and internal object structure. We present the first debate-driven agentic approach to articulated 3D object reconstruction from text or image inputs that both grounds articulation reasoning in concrete motion and exposes the occluded geometry revealed under articulation. High-level agents reason about object semantics and motion using knowledge from vision-language and video models, while low-level agents estimate articulation parameters and interaction points; together, they engage in a two-round structured debate that first exploits global--local disagreement and then grounds the agents in freely generated video. The same video prior, conditioned on the agreed articulation, then drives each part through its motion to expose occluded interiors and geometry that cannot be inferred from a single static view. By combining agentic reasoning with a video generative prior, our approach jointly infers articulation and reconstructs complete 3D articulated objects, producing high-fidelity geometry, internal structure, and motion-consistent states beyond directly observed surfaces.
関連論文
- MV-dVRK: 空間的外科知覚のための多視点ベンチマーク3D再構成
- 大規模再構成モデルを用いた人と物体のインタラクション再構成3D再構成
- PIVOT: 実世界3D再構成における姿勢・内部パラメータ・新視点評価のためのマルチ軌道データセットとテストベッド3D再構成
- OccamView: フレーム予算制約下のアクティブ3Dガウス再構成のためのオブジェクト条件付き視点選択3D再構成
- DerainSplat: スパースな雨天視点からのフィードフォワードによるクリーンな3Dガウススプラッティング3D再構成
- Stipple: 視覚慣性トラッキングによるリアルタイムインクリメンタルガウシアンスプラッティング3D再構成