SMART: 大規模合成事前学習によるゼロショットSim-to-Real関節物体マニピュレーション
SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
関節物体操作のための関節構造を考慮したシミュレーション基盤と分散合成システムを構築し、100万件超のデモデータでVLAモデルを事前学習することで、実機へのゼロショット転移を実現した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang, Shiqiang Zhu, Chenjia Bai, Xuelong Li
分類: cs.RO, cs.AI
原文アブストラクト
The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
関連論文
- EmbodiedSmith:シミュレーションにおける再帰的自己改善フライホイールによる身体性データのスケーリングsim2real
- 安全なリアルタイムロボット制御のためのマイクロニューラルポリシーsim2real
- 脚から車輪へ:移動ベースヒューマノイドのための身体性を考慮した人間動作リターゲティングsim2real
- ロボットはその記述ではない:形態認識ポリシーの表現堅牢性を評価するGaugeBenchsim2real
- UWB×Crazyflow:空中ロボティクス向け劣化フィードバックの大規模シミュレーションsim2real
- ROS 2上の無線対応ロボットナビゲーションのためのSionna-Isaac Sim閉ループ連成シミュレーションフレームワークのデモsim2real