日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
シーン表現arXiv:2610.02322

SCION: インスタンス化ニューラルプリミティブによるシーン構成

SCION: Scene Composition with Instanced Neural Primitives

シェア:XThreadsFacebookLINEはてブBluesky

再利用可能なプリミティブとその配置インスタンスでシーンを階層的に表現し、少ないパラメータで高品質な3D再構成と編集を可能にする手法。

詳しい要約

1. どんなもの?

多視点画像から3Dシーンを再構成する際、独立した数百万個のGaussianをそのまま保持するのではなく、再利用可能な少数のprimitive(語彙)と、それをシーン内に配置する軽量なworld-space instanceに置き換える階層的・構成的シーン表現SCIONを提案する。brickや草、小石、葉など実世界に繰り返し現れる要素を発見し、compactで操作可能な表現として保持する。1.2 MBでも高品質を維持し、instance単位の編集やアニメーションを再学習なしで可能にする。

2. 先行研究と比べてどこがすごい?

既存の3D Gaussian Splattingやその抽象化・圧縮手法は各要素を独立・ユニークとして扱い、シーンごとに数百万のGaussianをフィットし冗長なパラメータを保存し、下流タスクの操作ハンドルが弱い。Splat and Replaceなどはテンプレート物体をフィットするが、繰り返し要素の選択がほぼ手動。SCIONはprimitiveの語彙とinstance配置を自動的に発見し、rate-distortionで既存Gaussian圧縮法より有利で、再学習なしのinstance編集・アニメーションを可能にする点が異なる。

3. 技術・手法の肝は?

階層的構成的シーン表現として、独立Gaussianを再利用可能なprimitiveのcompactな語彙と、変換コピーをシーン全体に配置する軽量world-space instanceに置き換える。多視点キャプチャに対し、離散・連続のシーンパラメータをjoint optimizationでフィットする。splatとinstanceの2レベルdensificationと、共有primitive間でディテールを保つadversarial lossを組み合わせる。

4. どうやって有効だと検証した?

多視点キャプチャにフィットさせ、1.2 MBでも高品質を維持することを示す。rate-distortionが既存Gaussian圧縮法より有利であること、instance単位のシーン編集とアニメーションが再学習なしで可能であることを結果として示している。

5. 議論はある?

要旨からは不明。ただし主張として、ニューラルシーン表現はシーンを独立primitiveとして記憶する必要はなく、再利用可能な部品を発見できることを示すと述べている。

6. 次に読むべき論文は?

Splat and Replace、3D Gaussian Splatting、およびfollow-upのabstraction/compression手法が参照・比較されている。関連としてGaussian圧縮や階層的シーン表現の研究を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: William Koch, Amogh Joshi, Cyrus Vachha, Cheng Zheng, Felix Heide

分類: cs.CV

原文アブストラクト

Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hier- archical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this represen- tation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives. The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts. Project webpage: https://light.princeton.edu/SCION

関連論文

PR本紙発行元 EmplifAI