日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34550

視線プロンプト:Vision-Language-Actionファインチューニングのための時間的に密な人間の注意

Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

シェア:XThreadsFacebookLINEはてブBluesky

VRテレオペレーション中の視線をフレーム単位の視覚プロンプトとしてVLAモデルに与え、展開時は軽量予測器で視線を推定することで、実世界の両手マニピュレーション成功率を26.3%から56.0%に向上させた。

著者: Yihan Zhou, Rui Yan, Mingcong Li, Zheyuan Huang, Xu Yang, Xueyang Guo, Yilin Mo

分類: cs.RO

原文アブストラクト

Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $π_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.

関連論文

PR本紙発行元 EmplifAI