生成的なReal-to-Simによるロボットマニピュレーションのためのテスト時空間推論
Test-Time Spatial Reasoning for Robot Manipulation Using Generative Real-to-Sim
単一のRGB-D画像から3D生成モデルと視覚言語モデルでシミュレーション可能なシーンを再構築し、数千の並列物理シミュレーションと進化的探索によって物体配置を最適化する、訓練不要のテスト時空間推論フレームワークを提案。
著者: Ivan Kapelyukh, Yafei Hu, Ran Gong, Brandon May, Tushar Kusnur, Laura Herlant, Karl Schmeckpeper, Edward Johns, Xiaohan Zhang
分類: cs.RO, cs.CV
原文アブストラクト
Spatial reasoning is fundamental to general robot intelligence, as it enables robots to complete long-horizon tasks involving multi-object interaction. We introduce Simify, a training-free, test-time framework that performs explicit spatial reasoning via massively parallel physics simulation. From a single RGB-D image of a scene, Simify reconstructs simulation-ready assets leveraging 3D generative models and vision-language models. Then given a task specified by a reward function (e.g., build the tallest tower), Simify launches thousands of parallel rollouts in simulation and performs an evolutionary search to optimize object arrangements, typically converging within seconds. We conduct quantitative experiments on real-robot hardware to demonstrate the ability of our framework to execute complex object rearrangement tasks end-to-end with previously unseen objects. Results show that our framework outperforms prior work on foundation models for spatial reasoning by effectively exploiting large-scale parallel simulation during inference, and also highlight the importance of complete and accurate geometry for successful sim-to-real transfer.