日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.21112

単一スキャンからのガウシアンスプラッティングによる実演合成と視覚運動ポリシー学習

Demonstration Synthesis from a Single Scan via Gaussian Splatting for Visuomotor Policy Learning

シェア:XThreadsFacebookLINEはてブBluesky

1回の動画スキャンから3Dガウシアンスプラッティングで環境を再構成し、物理シミュレータなしで把持・軌道を生成して高忠実な実演を量産、拡散ポリシー学習に活用するデータエンジンを提案。

詳しい要約

1. どんなもの?

- 単一のビデオスキャンから高忠実度のデモンストレーションを大量生成するデータエンジン GaussianFactory を提案 - シーンを編集可能な 3D Gaussian Splatting (3DGS) レプリカとして再構成し、物体組み合わせタスクをサンプリング - 各タスクで把持と軌道を純粋に運動学的に計画し、写実的なデモをレンダリング - 物理エンジンを生成ループに使わず、接触力が結果を決める把持形成時のみ学習済み接触モデルを利用 - 合成デモのみで訓練した diffusion policy がシミュレーションで 95.1%、実機 UR10e で 84.2% の成功率を達成

2. 先行研究と比べてどこがすごい?

- 既存のデモ合成手法は手作業コスト、視覚的忠実度の限界、物理シミュレータへの重い依存という制約があった - GaussianFactory は人間入力が単一ビデオスキャンのみで、生成ループに物理エンジンを必要としない - 3DGS による編集可能な再構成と写実的レンダリングにより、目標環境に視覚的に一致するデモを生成 - 物理ダイナミクスは把持形成時の接触モデルのみに限定し、計算コストと複雑さを大幅に削減

3. 技術・手法の肝は?

- 単一ビデオスキャンからシーンを 3D Gaussian Splatting (3DGS) レプリカとして再構成 - 再構成から物体組み合わせタスクをサンプリングし、各タスクで把持と軌道を運動学的に計画 - 把持形成時の接触力相互作用のみ、事前学習済みの学習ベース接触モデルで処理 - 接触モデルはインタラクションデータセットで一度だけ事前学習 - 生成されたデモは写実的にレンダリングされ、diffusion policy の訓練に直接用いられる

4. どうやって有効だと検証した?

- シミュレーション環境と実世界ワークスペース(物理 UR10e ロボット)の 2 つのセットアップで end-to-end の scan-to-deployment ワークフローを実装 - シミュレーションは再現性のために実世界の代役として使用 - 合成デモのみで訓練した標準的な diffusion policy が、シミュレーションで 95.1%、実機で 84.2% の成功率を達成 - これにより合成デモの下流タスクへの有用性を検証

5. 議論はある?

- 要旨からは不明(限界や議論の詳細には触れられていない) - 物理エンジンを使わないアプローチの一般性や、接触モデルの適用範囲については言及なし - 実世界での成功率 84.2% とシミュレーションの 95.1% の差の要因分析は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として、既存の demonstration synthesis 手法、3D Gaussian Splatting (3DGS)、diffusion policy、学習ベース接触モデルが挙げられる - 同分野の定番として、模倣学習におけるデータ拡張や sim-to-real 転移に関する研究が次に読むべき候補

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Beichen Wang, Yuen-Hei Yeung, V. R. Sridhar Devarakonda, Xuesu Xiao

分類: cs.RO

原文アブストラクト

Training a visuomotor policy calls for abundant demonstrations that closely match the target environment, yet collecting them anew remains expensive. Existing demonstration synthesis methods reduce this cost but remain constrained by high manual effort, limited visual fidelity, or heavy reliance on physics simulators. This paper introduces GaussianFactory, a high-fidelity data engine that mass-produces demonstrations with a single video scan as its only human input and no physics engine in the generation loop. Specifically, GaussianFactory reconstructs the scene as an editable 3D Gaussian Splatting (3DGS) replica and samples from the object-combination tasks the scene affords. For each task, it plans grasps and trajectories purely kinematically on the geometric reconstruction, rendering photorealistic demonstrations that visually match the target environment. Physical dynamics enter the pipeline only where contact force interactions dictate the outcome---during grasp formation, via a learned contact model pretrained once on an interaction dataset. To evaluate the downstream utility of the synthesized demonstrations, we implement an end-to-end scan-to-deployment workflow in two setups: a simulated scene that stands in for the real world to enable reproducibility, and a real-world workspace with a physical UR10e robot. In each setup, a standard diffusion policy trained solely on the synthesized demonstrations achieves 95.1% and 84.2% success rates, respectively.

関連論文

PR本紙発行元 EmplifAI