VLAモデルの内部表現における自然状態予測精度が制御された応答性の弱さを隠す可能性
Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts
VLAモデルの内部表現から物体状態を予測する際、自然な観測での高精度が必ずしも制御された状態変化への応答性を意味しないことを示し、予測精度・応答性・文脈安定性を分離して評価する枠組みを提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Hyungjoon Kim, Wonbin Son, Mi Young Lee, Jun Young Lee, Seungmin Rho
分類: cs.RO, cs.CV, cs.LG
原文アブストラクト
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing contextual cues to contribute to prediction. In this paper, we introduce an evaluation framework that separates prediction accuracy, target-state responsiveness, and context stability using physically validated observations that cross target coordinates with robot contexts. We demonstrate that high natural-trajectory accuracy can coexist with weak controlled target-state responsiveness in fixed representation-readout pairs. Comparisons and interventions involving representations, readouts, and training data show that the three properties provide distinct diagnostic information. Furthermore, adding responsiveness and context sensitivity to a failure predictor based on initial state error and physical variables reduces policy-failure prediction error on new initializations relative to the specified baseline while same-observation controlled MAE is also informative. These findings motivate evaluating target-state responsiveness and context stability alongside natural prediction accuracy, and examining their relationship to actual policy behavior and task outcomes.