日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11248

SimVLA:モバイルマニピュレーションのためのゼロショットSim-to-Real VLA学習

SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーションのみでVLAを学習し、実世界のモバイルマニピュレーションへゼロショット転移するフレームワークSimVLAを提案。実機デモ50件で学習した方策を上回る性能を示した。

詳しい要約

1. どんなもの?

- モバイルマニピュレーションのためのVLA(Vision-Language-Action)モデルを、シミュレーションのみで学習するフレームワーク「SimVLA」を提案。 - テレオペレーションなしで、合成シミュレーションデータのみを用いてエンドツーエンドで学習。 - 実世界の家庭環境を含むゼロショットsim-to-real転移を実現。

2. 先行研究と比べてどこがすごい?

- 従来のVLAは実世界データ収集のコストと複雑さに制限されていた。 - シミュレーションはスケーラブルな代替手段だが、モバイルマニピュレーションにおけるsim-to-real VLA学習の可能性はほとんど未探索だった。 - SimVLAは、実世界の50件のドメイン内デモンストレーションで訓練されたポリシーを上回る性能を示し、シミュレーションがスケーラブルなsim-to-realモバイルマニピュレーションを可能にすることを示唆。

3. 技術・手法の肝は?

- 2つの補完的なシミュレーション由来データセットで事前学習:SimAction(35の多様なモバイルマニピュレーションタスクにわたる大規模ロボット行動データセット、原子スキルの合成により生成)とSimVQA(特権シミュレータ状態を活用し、空間的・幾何学的・サブタスクレベルの視覚言語監督を提供)。 - その後、SimActionとSimDeploy(多様なシミュレーション環境でのポリシーロールアウトから収集されたデータセット)の混合でポストトレーニング。 - 複数の補完的な監督形式の価値を実証。

4. どうやって有効だと検証した?

- 棚卸し、注ぎ、掃除などのタスクで評価。 - 実世界の家庭環境を含むゼロショット転移を実証。 - 実世界の50件のドメイン内デモンストレーションで訓練されたポリシーを上回る性能を確認。

5. 議論はある?

- シミュレーションがスケーラブルなsim-to-realモバイルマニピュレーションを可能にすることを示唆。 - VLAトレーニングにおいてシミュレーションを効果的に活用するための、複数の補完的な監督形式の価値を実証。 - 具体的な議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:実世界の50件のドメイン内デモンストレーションで訓練されたポリシー。 - 関連手法:VLA(Vision-Language-Action)モデル、sim-to-real転移、モバイルマニピュレーション。 - 同分野の定番:RT-1、RT-2、PaLM-E、OpenVLAなどが考えられるが、要旨では明示されていない。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kyoungin Baik, Youngwoon Lee

分類: cs.RO

原文アブストラクト

Large-scale, diverse datasets have driven the success of LLMs and VLMs. But VLAs for robotics remain limited by the cost and complexity of real-world data collection. While simulation offers a scalable alternative, its potential for sim-to-real VLA learning in mobile manipulation remains largely underexplored. We introduce SimVLA, an end-to-end framework that trains VLAs entirely on synthetic simulation data without teleoperation for mobile manipulation. SimVLA is first pre-trained on two complementary simulation-derived datasets: SimAction, a large-scale robot action dataset spanning 35 diverse mobile manipulation tasks, generated by composing atomic skills, and SimVQA, which leverages privileged simulator state to provide spatial, geometric, and subtask-level visual-language supervision. We further post-train SimVLA on a mixture of SimAction and SimDeploy, a dataset collected from policy rollouts across diverse simulated environments. We evaluate SimVLA on tasks including restocking, pouring, and cleaning, and show zero-shot transfer to real-world mobile manipulation, including real home environments. SimVLA outperforms policies trained on 50 in-domain real-world demonstrations, suggesting that simulation can enable scalable sim-to-real mobile manipulation. We further demonstrate the value of multiple complementary forms of supervision for effectively leveraging simulation in VLA training.

関連論文

PR本紙発行元 EmplifAI