日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30868

VLaRL: シミュレーション学習の潜在条件付き残差RLによる視覚言語行動モデルの拡張

VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL

シェア:XThreadsFacebookLINEはてブBluesky

凍結したVLAモデルに残差強化学習を適用し、シミュレーションで学習した制御を実機に転移する手法を提案。VLAの内部潜在表現を条件とすることで、実機での追加学習なしに接触の多い操作タスクの成功率を向上させた。

詳しい要約

1. どんなもの?

- VLaRLは、Vision-Language-Action (VLA) モデルの物理実行精度を向上させる手法。 - 凍結したVLAに対し、残差強化学習 (residual RL) をシミュレーションで訓練し、実機に展開する。 - 実世界RLやオンライン適応を必要とせず、接触の多いマニピュレーションタスクで成功を改善。

2. 先行研究と比べてどこがすごい?

- 従来の残差RLは実機での訓練が高コスト・安全面で問題。 - シミュレーション訓練した残差ポリシーを実機に転移する際の視覚ギャップが課題。 - VLaRLはピクセルレベルの視覚対応を必要とせず、VLAの内部潜在表現を条件と転移インタフェースに使用。 - 軽量マッパーでシミュレーション由来の潜在を実機の潜在分布に変換。

3. 技術・手法の肝は?

- VLAの内部vision-language潜在表現を残差制御の条件付けに使用。 - シミュレーションと実機のsim-to-real転移インタフェースとしても同潜在を利用。 - シミュレーションで訓練した残差ポリシーを実機に展開するため、軽量マッパーを学習し、シミュレーション潜在を実機潜在分布に変換。 - 実世界RLやオンライン適応は不要。

4. どうやって有効だと検証した?

- 4つの接触の多いマニピュレーションタスクと2つのVLAバックボーンで評価。 - すべてのタスク・バックボーンの組み合わせで実世界成功率を改善。 - 制御されたアブレーションにより、潜在条件付けと潜在アラインメントの両方の重要性を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、Residual RL、Vision-Language-Action (VLA) モデル、Sim-to-Real Transfer、Latent Representation Alignment などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Namiko Saito, Kinam Kim, Heecheol Kim, Katsushi Ikeuchi, Yasuyuki Matsushita

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) models provide broad, instruction-conditioned manipulation behaviors, but their physical execution can remain imprecise during contact-rich interaction. Residual reinforcement learning (RL) can correct such errors while keeping the VLA frozen, but real-robot RL is costly and safety-critical. We propose VLA Latent-Conditioned RL (VLaRL), which enables residual RL for frozen VLAs to be trained in simulation and deployed on real robots without real-world RL or online adaptation. The key challenge is transferring the learned residual policy despite the visual gap between simulation and reality. Rather than requiring pixel-level visual correspondence, VLaRL uses the VLA's internal vision-language latent representation to condition residual control and as the sim-to-real transfer interface, and learns a lightweight mapper that transforms simulation-derived latents toward the real latent distribution. Across four contact-rich manipulation tasks and two VLA backbones, VLaRL improves real-world success in all task-backbone combinations, while controlled ablations demonstrate the importance of both latent conditioning and latent alignment for transferring simulation-trained residual control.

関連論文

PR本紙発行元 EmplifAI