日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D再構成arXiv:2610.11215

GATOR: 日常画像からの生成的・エージェント的3D物体再構成

GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images

シェア:XThreadsFacebookLINEはてブBluesky

1枚以上の日常的な画像から、遮蔽部分を推論しつつテクスチャ付き3D物体とシーン内での姿勢を復元する生成的・エージェント的フレームワークを提案。

詳しい要約

1. どんなもの?

- カジュアル画像から完全な3Dオブジェクトを再構成するGATORを提案。 - 生成とエージェントを組み合わせたフレームワーク。 - テクスチャ付きオブジェクト資産とシーン相対poseを1枚以上の画像から復元。 - 対象は合成物体、散らかったテーブル、屋内シーン。

2. 先行研究と比べてどこがすごい?

- 先行研究との具体的比較は要旨からは不明。 - 疎で不確実な観測の統合と遮蔽された表面の推論を扱う点を強調。 - 生成資産をエージェント初期化に用い、編集-描画-レビューのループで改良する点が特徴。

3. 技術・手法の肝は?

- local modality mixerがpatch-aligned RGB、target-mask、pointmap特徴を結合し、cross-view reasoning前にシーン文脈を保持。 - text-guided semantic conditioningがカテゴリ名とキャプションをstage-specific adaptersで構造・幾何・外観生成に補完。 - 生成資産がmultimodal agentを初期化し、observation-guided edit-render-review loopで構造とテクスチャを改良。

4. どうやって有効だと検証した?

- 合成物体、散らかったテーブル、屋内シーンで評価。 - 幾何と外観の忠実性、疎観測からのシーン相対pose復元を確認。 - time-budget比較とシーンレベルシミュレーションで再構成効率とシミュレーション準備性を示す。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗ケース、計算コストの詳細な議論は記載されていない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として3D reconstruction、pose estimation、generative models、multimodal agentsの定番研究を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Qirui Wu, Stan Birchfield, Hesam Rabeti, Angel X. Chang, Bowen Wen

分類: cs.CV, cs.RO

原文アブストラクト

Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: https://research.nvidia.com/labs/lpr/gator/

関連論文

PR本紙発行元 EmplifAI