日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
arXiv:2608.04246

SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

シェア:XThreadsFacebookLINEはてブBluesky

著者: Harshitha Rajaprakash, Aditeya Prajapati, Rong Xue, Abrar Anwar, Jesse Thomason

分類: cs.RO, cs.CV

原文アブストラクト

Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.