日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3D表現学習arXiv:2609.19463

ParticleSplat: 自己教師ありオブジェクト中心潜在粒子スプラッティング

ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

シェア:XThreadsFacebookLINEはてブBluesky

複数視点の画像から3Dガウシアンスプラッティングを用いてシーンを物体ごとの潜在粒子に分解し、教師なしで物体マスクを獲得しロボット操作タスクの性能を向上させる手法を提案。

詳しい要約

1. どんなもの?

- 自己教師あり物体中心学習手法 ParticleSplat を提案 - シーンを意味的実体を表す潜在粒子に分解 - feedforward 3D Gaussian Splatting を利用 - Deep Latent Particles (DLP) を基盤に拡張 - 3D 潜在粒子空間を新たに導入 - 新規視点合成目的で学習 - 複数視点とカメラポーズを共有3D物体中心潜在空間に符号化 - 粒子を particle-aligned 3D Gaussians に変換しシーン再構成

2. 先行研究と比べてどこがすごい?

- DLP の2D性を克服し3D空間・幾何推論を可能に - ロボット操作など下流タスクに重要な3D推論を実現 - 潜在粒子と3D Gaussian primitives の構造類似性を活用 - 教師なしで物体マスクを学習 - 潜在空間での粒子操作による制御可能な3Dシーン編集を実現 - 学習された3D表現がロボット操作タスクの性能を向上

3. 技術・手法の肝は?

- 自己教師あり物体中心学習 - feedforward 3D Gaussian Splatting を採用 - 複数視点とカメラポーズを共有3D物体中心潜在空間に符号化 - 粒子を particle-aligned 3D Gaussians に変換 - 新規視点合成目的で3D潜在粒子空間を学習 - 粒子の属性(位置、スケール、外観など)を保持

4. どうやって有効だと検証した?

- シミュレーションおよび実世界データセットで評価 - 教師なしで物体マスクを学習することを確認 - 潜在空間での粒子操作による3Dシーン編集を実証 - ロボット操作タスクの下流性能向上を確認

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- Deep Latent Particles (DLP) - 3D Gaussian Splatting - 物体中心学習 - ロボット操作

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lyuxing He, Daniel Guo, Elizabeth Terveen, Deepak Pathak, David Held, Tal Daniel

分類: cs.CV, cs.RO

原文アブストラクト

We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which represents images as a set of particles with attributes such as position, scale, and visual appearance, we address a key limitation of DLP: its inherently 2D nature, which prevents explicit 3D spatial and geometric reasoning that are critical for downstream tasks such as robotic manipulation. Leveraging the structural similarity between latent particles and 3D Gaussian primitives, we introduce a 3D latent particle space trained with a novel view synthesis objective. Our model jointly encodes multiple views with camera poses into a shared 3D object-centric latent space, then transforms particles into particle-aligned 3D Gaussians whose composition reconstructs the full scene. On simulated and real-world datasets, we show that this formulation inherently learns object masks without supervision and supports controllable 3D scene editing, such as moving objects by modifying particles in the latent space. We further establish that the learned 3D representation improves downstream performance on robotic manipulation tasks.

関連論文

PR本紙発行元 EmplifAI