日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.02804

SARI: 接触の多いマニピュレーションのためのフェーズ分割Sim-Real協調訓練

SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

自由空間の接近はシミュレーションで多様に生成し、接触物理は実機で少数収集するフェーズ分割協調訓練により、未知配置でも成功するVLAポリシーを実現した。

詳しい要約

1. どんなもの?

- 視覚言語行動(VLA)モデルを用いた接触を伴う操作タスクの適応を目的としたフレームワーク。 - SARI(Simulated Approach, Real Interaction)と名付けられた、シミュレーションと実世界の共訓練手法。 - フェーズを分割し、自由空間のアプローチはシミュレーションで、接触インタラクションは実世界で学習する。 - 実世界のデモンストレーション収集コストを削減しつつ、未知の配置への汎化を実現する。

2. 先行研究と比べてどこがすごい?

- 従来のVLAモデルは、接触を伴う操作タスクへの適応に高コストな実世界デモを必要とした。 - 完全なシミュレーションと実世界の共訓練ベースラインは、未知の配置で完全に失敗(成功率0%)するのに対し、SARIは27.5%の成功率を達成。 - 実データ収集時間を34.3%削減。 - フェーズ分割により、各ドメインの強みを活かす点が新しい。

3. 技術・手法の肝は?

- 自由空間のアプローチはシミュレーションで多様に生成し、接触インタラクションは実世界で少数の配置のみ収集。 - フォトリアルなデジタルツインを用いてシミュレーションアプローチを生成。 - フェーズ分割されたデモンストレーションでポストトレーニングし、単一のポリシーがシミュレーションと実世界の接触をシームレスに統合。 - 視覚的外観のアラインメントと共有カメラ相対行動表現を使用し、明示的なフェーズラベルや手動スイッチは不要。

4. どうやって有効だと検証した?

- 5つの実世界の接触を伴う操作タスクで評価。 - 実データ収集時間を34.3%削減。 - 未知の配置での成功率27.5%を達成し、完全なシミュレーションと実世界の共訓練ベースライン(成功率0%)を上回った。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。関連手法として、VLAモデル、sim-real co-training、contact-rich manipulationの定番研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xingxin He, Yuxuan Jiang, Haonan Zhang, Chuhan Cui, Kaile Li, Zhongxing Zheng, Caihao Xu, Ziqi Wang

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) models often require costly real-world demonstrations to adapt to contact-rich manipulation tasks, particularly when generalization across object placements is needed. We propose SARI (Simulated Approach, Real Interaction), a phase-split sim-and-real co-training framework built on a simple insight: spatial coverage and contact physics should be acquired from the domains best suited to them. Specifically, free-space approaches require spatial diversity but tolerate modest simulation gaps, making them ideal for synthetic generation; conversely, contact interactions demand accurate physics but vary little across object placements, allowing a few real demonstrations to generalize across the workspace. SARI generates diverse simulated approaches in a photorealistic digital twin while collecting real contact interactions at only a few placements. Post-trained on these phase-segmented demonstrations, a single policy seamlessly stitches simulated approaches with real contact interactions using visual appearance alignment and a shared camera-relative action representation--without explicit phase labels or hand-coded switches. Across five real-world contact-rich manipulation tasks, SARI reduces real-data collection time by 34.3% and achieves 27.5% success at unseen placements, where all full-task sim-real baselines fail completely (0%).

関連論文

PR本紙発行元 EmplifAI