VLA事後学習におけるデータ拡張の正しい使い方と落とし穴
How (and How Not) to Use Data Augmentation in VLA Post-Training
VLAモデルの強化学習による事後学習で、画像拡張はcriticにのみ適用しactorには適用しないことが重要だと示し、分布外タスクの成功率を改善する実践的な指針を提示した。
著者: Bram Grooten, Joaquin Vanschoren
分類: cs.RO, cs.LG
原文アブストラクト
Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with reinforcement learning (RL) has been shown to benefit robustness, but significant room for improvement remains. In this work, we systematically study the effect of image augmentation on VLA post-training. We find that it is crucial to augment only the critic module during RL updates, while leaving the actor's input clean during both rollouts and updates. For $π_{0.5}$ and GR00T N1.5 this raises out-of-distribution success on LIBERO-Plus by $7.8$ and $10.0$ points respectively, while augmenting the actor collapses training entirely. We investigate a range of augmentation types and strengths, and provide practical recommendations for improving generalization in VLA post-training.