R2S-EGO: スパースキャプチャ実世界からシミュレーションへのデュアルプロキシ精緻化
R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim
ロボットの軌跡に沿った視点を効率的に補完するため、シミュレータ由来のロボットプロキシとキャプチャ由来の幾何プロキシを組み合わせ、実画像をアンカーに疑似観測を生成してシーンを精緻化する手法を提案。
著者: Shuai Fang, Xin Deng, Yuchen Kang, Zhenjiang Li, Jie Chen
分類: cs.RO, cs.CV, cs.GR
原文アブストラクト
Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet dense multi-view capture limits per-environment real-image capture-count efficiency, and sparse human capture can leave behavior-scoped robot views under-supported. Camera-controlled synthesis can fill missing views, but its use in R2S requires behavior-admissible queries and capture-anchored structural conditioning. We present R2S-EGO, which couples a simulator-derived robot proxy that represents the behavior-scoped executable query domain with a capture-anchored geometry proxy that supplies scene-specific structural conditions. Within this domain, fixed- budget selection targets current support deficits for which geometry support is available. The generated observations are assimilated as pseudo-observations to refine the visual asset, while real captures remain anchors. The fused geometry proxy also supplies the scene collision surface, which is refreshed between rounds. Together, these updates refine the existing simulation scene while its robot dynamics and control stack stay fixed. Across 48 frozen Unitree G1 ego views in three Replica scenes, six-view R2S-EGO reaches 19.062 dB PSNR, compared with 14.226 dB for the strongest reported R2S baseline. Across five paired policy-training seeds, R2S-EGO achieves 82.5% +/- 6.8% real-G1 sitting success, compared with 10.0% +/- 10.5% for GaussGym.