ガードが見逃すものをロボットは実行する:VLA指示における暗示的害
What the Guard Misses, the Robot Executes: Implied Harm in VLA Instructions
VLAモデルは拒否できず指示を実行するため、有害な意図の検知はモニタに依存する。タスクを固定し意図の明示度だけを変えて検証した結果、テキストガードは露骨な要求は検知するが暗示的な害の最大95%を見逃し、活性化監視でも補えないことを示した。
著者: Sripad Karne, Arjun Balaji
分類: cs.RO
原文アブストラクト
Vision-language-action models (VLAs) act on instructions without being able to refuse, so screening harmful requests falls to monitors. We test whether these monitors catch ordinary robot tasks requested for harmful reasons, holding the task fixed while varying only how explicitly the intent is stated. $π_{0.5}$ completes the task at every level of explicitness, as often as for harmless controls. Text guards flag nearly every blunt request but few implied ones: up to 95% of implied-harm runs end with the task done and no flag raised, and up to 90% even after recalibrating on robot instructions. Monitoring the model's activations does not close this gap. Linear probes separate harmful from harmless instructions almost perfectly in the base language and vision-language models, but this separation weakens after robot training in two model families, most for implied harm.