TimelyDAgger: VLAポリシー改善のためのタイミング考慮型エキスパートクエリ
TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement
VLAポリシーの内部特徴を監視し、エキスパートの介入タイミングを適応的に調整することで、限られた介入予算でポリシー改善を効率化する手法を提案。
著者: Zhixuan Zhao, Peiyan Li, Enhao Zhang, Yueran Tao, Hao Wang, Chenghao Yue, Lei Lv, Wentao Zhao, Jiahao Chen, Xin Liu, Kangyao Huang, Yu Luo, Huaping Liu
分類: cs.RO
原文アブストラクト
DAgger improves robot policies by aggregating expert supervision from states visited during policy execution. Robot-gated DAgger automates expert queries, allowing the robot to decide when to request expert takeover. While existing gates emphasize detecting the need for assistance, takeover timing also shapes the content of these demonstrations and their value for policy learning. We propose TimelyDAgger, combining Bridge-PCA monitoring of internal vision-language-action (VLA) features with Feedback-guided Threshold Adaptation based on expert behavior to improve takeover timing. We introduce an evaluation framework linking failure detection, takeover timing, and policy improvement, including Target-Aligned Supervision Ratio (TASR) for assessing supervision quality without retraining. Experiments show that takeover timing affects policy learning, with TimelyDAgger achieving competitive failure detection and higher post-training success in most evaluated settings under matched expert-action budgets.