Pelican-VLA 0.5: 行動前に注意を向けることで汎化を向上
Pelican-VLA 0.5: Attending Before Acting Benefits Generalization
視覚言語理解、未来フレーム生成、行動予測を統合したVLAモデルを提案し、注意レベルの汎化を実現。ボトルネックトークンにより操作対象への注意が向上することを示した。
著者: Zeyuan Ding, Wenhai Liu, Yang Xu, Jiayu Hu, Yinda Chen, Yi Zhang, Yong Dai, Jian Tang, Xiaozhu Ju
分類: cs.RO, cs.LG
原文アブストラクト
In this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, future-frame generation, and action prediction within a single architecture. Pelican-VLA 0.5 achieves attention-level generalization: without object annotations, segmentation masks, attention supervision, or task-specific fine-tuning, its action pathway already focuses on the manipulation-relevant object and contact region. This behavior persists across unseen scenes and unseen robot embodiments, and is substantially stronger than in other open-source VLA baselines. We verify that this ability originates from the learnable Bottleneck Token inserted between perception and action: by routing task-relevant visual information through a compact bottleneck, the tokens interface induces manipulation-centric attention during pre-training and remains effective across different policy structures, including a MoT-style architecture.