日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.07652

SMART: 大規模合成事前学習によるゼロショットSim-to-Real関節物体マニピュレーション

SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining

シェア:XThreadsFacebookLINEはてブBluesky

関節物体操作のための関節構造を考慮したシミュレーション基盤と分散合成システムを構築し、100万件超のデモデータでVLAモデルを事前学習することで、実機へのゼロショット転移を実現した。

詳しい要約

1. どんなもの?

- 本論文は、関節物体操作のための大規模合成デモンストレーションを活用したスケーラブルシステム SMART を提案する。 - 中核は SMART-Sim という関節認識設計のシミュレーションプラットフォームで、タスク生成とデモ収集を効率化する。 - これを用いて 44 の原子タスクタイプ、5 のロボットセットアップ、2,507 の関節物体にわたる 100 万件以上のデモからなる SMART-Data を合成する。 - SMART-Data で事前学習した vision-language-action (VLA) モデルは、シミュレーションベンチマークで競争力のある性能を示し、実世界の関節物体操作タスクで zero-shot sim-to-real 転移とスケーラブルな性能を達成する。

2. 先行研究と比べてどこがすごい?

- 既存の合成データ取り組みは限られた関節物体カテゴリしかカバーしていなかった。 - 汎用合成パイプラインは、パーツレベルの意味論や関節制約に対する明示的な設計を欠いており、エージェント的タスク生成や高品質な関節操作デモのスケーラブル合成を妨げていた。 - 本研究は、関節認識設計の SMART-Sim とエージェント的タスク生成、分散合成システムを組み合わせ、大規模で多様なデモを合成可能にした点が先行研究と比べて優れている。

3. 技術・手法の肝は?

- SMART-Sim は関節認識設計を備えたシミュレーションプラットフォームで、効果的なタスク生成と効率的なデモ収集を可能にする。 - エージェント的タスク生成を適用し、スケーラブルな分散合成システムを設計する。 - これらを用いて SMART-Data を合成する。SMART-Data は 44 の原子タスクタイプ、5 のロボットセットアップ、2,507 の関節物体にわたる 100 万件以上のデモを含む。 - このデータで vision-language-action (VLA) モデルを事前学習する。

4. どうやって有効だと検証した?

- SMART-Data で事前学習した VLA モデルがシミュレーションベンチマークで競争力のある性能を示すことを確認した。 - 実世界の関節物体操作タスクにおいて zero-shot sim-to-real 転移を達成し、スケーラブルな性能を示すことを検証した。 - これにより、合成デモが接触リッチな関節物体操作における VLA モデル性能向上のための効果的かつスケーラブルな監督を提供する可能性を実証した。

5. 議論はある?

- 合成デモンストレーションが、接触リッチな関節物体操作における VLA モデルの性能を改善するための効果的かつスケーラブルな監督を提供する可能性を強調している。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、vision-language-action (VLA) モデル、sim-to-real 転移、関節物体操作、合成データ生成パイプラインが挙げられる。 - 同分野の定番として、ロボット操作のための大規模事前学習やシミュレーションからの転移学習に関する研究が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang, Shiqiang Zhu, Chenjia Bai, Xuelong Li

分類: cs.RO, cs.AI

原文アブストラクト

The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.

関連論文

PR本紙発行元 EmplifAI