GIF: ロボット学習のための対話的・機能的オブジェクト構成のエージェント生成
GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning
ロボット操作の基盤モデル学習用に、対話的で機能的なオブジェクト構成を自動生成するエージェントフレームワークGIFを提案。2D/3D生成モデルとVLM検証を組み合わせ、衝突率1%未満で高品質なシーン生成を実現し、シミュレーションと実世界でのポリシー学習に有効性を示した。
著者: Long Xu, Zhiqi Zhang, Mi Yan, Shengliang Deng, Chong Xia, Mingyu Dong, Jiayi Chen, Jiangran Lyu, Fei Gao, Zhizheng Zhang, He Wang
分類: cs.RO
原文アブストラクト
Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios, with simulation providing an environment for both. Automated scene generation offers a promising path, yet prior work has largely emphasized coarse-grained scene layouts rather than fine-grained functional object compositions. Motivated by this gap, we present GIF, an agentic Generation framework for Interactive and Functional object compositions. In this framework, we recast this problem as disentangled reconstruction followed by relative pose recovery. CoGen produces instance-disentangled meshes with coarse initial poses leveraging complementary strengths of 2D and 3D generative models. GPRM refines the relative pose under joint geometric and physical guidance, and a VLM verifier selects the candidate that best matches the structured specification. We further construct a benchmark spanning eight representative contact-geometry classes and compare with state-of-the-art generators; GIF improves both asset quality and relation matching, while reducing collision rate to below 1%. Finally, we synthesize data for policy learning, revealing diversity scaling in both simulation and real-world deployment.