MaskVLA: 視覚マスキングによるVision-Language-Actionモデルの軌道過学習の抑制
MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model
VLAモデルの微調整時にメインカメラの視覚情報をランダムに一部マスクすることで、手首カメラ情報の活用を促し、軌道過学習を防いで汎化性能とタスク成功率を向上させる手法を提案。
著者: Yuxuan Jiang, Jiaying Huang, Ge Wang, Shenhao Yan, Jiahao Yang, Chengsi Yao, Qi Liu, Qing Zhao, Shuguang Cui, Yiming Zhao, Yatong Han, Zhen Li
分類: cs.CV, cs.AI, cs.RO
原文アブストラクト
Vision-Language-Action (VLA) models integrate vision-language understanding with executable robot actions, enabling end-to-end learning for robot control. However, our empirical analysis reveals that existing models exhibit severe trajectory overfitting when finetuned on limited datasets. To guide the model in effectively utilizing wrist camera information, we propose MaskVLA, a masking-based fine-tuning strategy. By randomly masking a small portion of the main camera's visual information, the model is guided to autonomously learn more fine-grained, task-relevant, and effective visual features. This process leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance. Our method has been comprehensively evaluated on RoboTwin 2.0, achieving an average success rate improvement of 23.2% and 16.8% compared to $π_0$ and OpenVLA-OFT, respectively. Furthermore, experiments on real-world ALOHA robots also demonstrate the effectiveness of our approach.