先行研究では、hidden-state-based risk probesとfunctional conformal predictionによる障害検出が提案されているが、その信頼性はキャリブレーションデータが展開条件と一致することに依存していた。SAFECASTは、contrast set perturbationsを導入することで、キャリブレーションデータが展開条件と異なる場合でも検出性能を向上させる点が新しい。また、視覚と言語の両方のperturbationsを用いることで、より効果的であることを示している。
3. 技術・手法の肝は?
SAFECASTの核心は、contrast set perturbationsを用いてhidden-state probeのトレーニングとキャリブレーションを強化することである。具体的には、視覚的(clutter、lighting changesなど)と言語的(reworded instructions)なperturbationsをデータに適用し、プローブが分布シフトに対してより敏感になるように訓練する。さらに、functional conformal predictionを用いて、検出結果の信頼性を調整する。
4. どうやって有効だと検証した?
実世界のDROIDとシミュレーションのLIBERO実験において、複数のVLMバックボーンを用いて、SAFECASTの障害検出ROC-AUCスコアが最先端のベースラインと比較して統計的に有意に向上することを検証した。また、視覚と言語の両方のcontrast set perturbationsを用いた場合に最も効果が高いこと、およびcontrast set perturbationsを用いるとsim-to-realキャリブレーションが実ロールアウトデータのみを用いるよりも優れたプローブを生成することを確認した。
5. 議論はある?
要旨からは、SAFECASTの限界や潜在的な欠点についての議論は不明である。ただし、contrast set perturbationsの生成方法や、異なるタイプのシフトに対する感度の違いなど、さらなる分析が必要かもしれない。また、実世界での適用における計算コストや、より複雑な環境での一般化については言及されていない。
Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.