日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.03498

検出して抑制:VLAモデルにおける敵対的パッチへのメカニズム的防御

Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルの内部表現をスパースオートエンコーダで解析し、敵対的パッチの存在と相関する特徴を検出して推論時に抑制することで、ファインチューニングなしにロバスト性を向上させる手法を提案。

詳しい要約

1. どんなもの?

- VLAモデルに対するadversarial patch攻撃の防御手法を提案。 - sparse autoencoder (SAE)で内部表現を解析し、攻撃と相関するfeatureを同定。 - 推論時にlinear probeで攻撃検出した場合のみfeatureを抑制。 - fine-tuning不要でrobustnessを向上。 - LIBERO-10で評価。

2. 先行研究と比べてどこがすごい?

- 従来のVLA adversarial defenseはfine-tuningや再学習が必要なものが多い。 - 本研究は内部機構の解析に基づき、推論時の条件付き介入で対処。 - fine-tuningのコストなしでrobustnessを改善。 - 攻撃関連の内部表現を標的とする点が新しい。

3. 技術・手法の肝は?

- sparse autoencoder (SAE)でVLAの内部表現を解析。 - adversarial patchの存在と強く相関するfeatureを同定。 - linear probeで攻撃を検出。 - 検出時のみ同定したfeatureを推論時に抑制。 - 継続的抑制ではなく条件付き介入が重要。

4. どうやって有効だと検証した?

- LIBERO-10上でVLA adversarial patch攻撃に対して評価。 - 条件付き介入は間欠的攻撃下でsuccess rateを改善。 - 継続的介入はpolicy performanceを大幅に低下させることを確認。 - 攻撃関連内部表現が防御標的として有用であることを示す。

5. 議論はある?

- 条件付き介入のタイミング制御がnominal policyへの影響を抑える上で重要。 - 攻撃関連内部表現の抑制が有効な防御戦略となり得る。 - 継続的介入は性能劣化を招くため、検出精度が鍵。 - 他の攻撃タイプやモデルへの一般化は要旨からは不明。

6. 次に読むべき論文は?

- VLAモデル(例: RT-2, OpenVLA)のadversarial robustnessに関する研究。 - sparse autoencoder (SAE)を用いた解釈可能性研究。 - adversarial patch攻撃の防御手法(例: adversarial training, input purification)。 - LIBERO-10ベンチマークを用いたロボット制御研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi

分類: cs.RO, cs.AI

原文アブストラクト

Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyze VLA representations using a sparse autoencoder (SAE) and identify a feature whose activation strongly correlates with the presence of an adversarial patch. Based on this analysis, we suppress the identified feature at inference time only when a linear probe detects an attack. This intervention improves robustness without the cost of fine-tuning the VLA. We evaluate our method against VLA adversarial patch attacks on LIBERO-10. Conditional intervention improves success rate under intermittent attacks, whereas continuously applying the same intervention substantially degrades policy performance. These results show that attack-related internal representations can provide useful targets for VLA adversarial defense and that controlling when to intervene is important for limiting disruption to nominal policy behavior.

関連論文

PR本紙発行元 EmplifAI