日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.23565

MaskVLA: 視覚マスキングによるVision-Language-Actionモデルの軌道過学習の抑制

MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルの微調整時にメインカメラの視覚情報をランダムに一部マスクすることで、手首カメラ情報の活用を促し、軌道過学習を防いで汎化性能とタスク成功率を向上させる手法を提案。

著者: Yuxuan Jiang, Jiaying Huang, Ge Wang, Shenhao Yan, Jiahao Yang, Chengsi Yao, Qi Liu, Qing Zhao, Shuguang Cui, Yiming Zhao, Yatong Han, Zhen Li

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models integrate vision-language understanding with executable robot actions, enabling end-to-end learning for robot control. However, our empirical analysis reveals that existing models exhibit severe trajectory overfitting when finetuned on limited datasets. To guide the model in effectively utilizing wrist camera information, we propose MaskVLA, a masking-based fine-tuning strategy. By randomly masking a small portion of the main camera's visual information, the model is guided to autonomously learn more fine-grained, task-relevant, and effective visual features. This process leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance. Our method has been comprehensively evaluated on RoboTwin 2.0, achieving an average success rate improvement of 23.2% and 16.8% compared to $π_0$ and OpenVLA-OFT, respectively. Furthermore, experiments on real-world ALOHA robots also demonstrate the effectiveness of our approach.

関連論文

PR本紙発行元 EmplifAI