ビットフリップ攻撃による視覚言語行動モデルの脆弱性:行動デコードアーキテクチャが攻撃耐性を左右する
Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability
量子化された視覚言語行動(VLA)モデルに対する初のビットフリップ攻撃を提案し、少数のフリップで閉ループ成功率を0%にできることを示した。攻撃に必要なフリップ数は行動ヘッドの種類に依存し、直接回帰やトークンポリシーでは1〜5回、フローマッチングでは100〜300回必要である。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yudong Gao, Linghan Chen, Wenhan Wu, Mia Zhou, Jiyao Wang, Kaiyan Ji, Mingyu Guo, Honglong Chen
分類: cs.CR, cs.AI
原文アブストラクト
Quantized Vision-Language-Action (VLA) models expose a weight-fault surface: Rowhammer-style faults can corrupt deployed INT8 bits. We present the first bit-flip attack on a VLA: a few gradient-selected flips reduce closed-loop success to $0\%$, while hundreds of random flips are harmless. Across four model variants spanning three action-head families, damaging bits concentrate in a few action-generating layers, but the empirical budget depends sharply on the head: direct regression and token policies fall in $1$--$5$ flips, whereas the evaluated flow-matching policies require ${\sim}100$--$300$. Our fixed-direction manifold-escape loss cuts \pizero{}'s budget from ${\sim}1000$ to ${\sim}100$ flips, and a matched five-direction sweep shows that the attack is not specific to an all-positive direction. On a direct head, protecting $3.1\%$ of weights preserves $60\%$ success at $K{=}100$, and protecting $5.3\%$ moves the open-loop break threshold from 3 to 100 flips. Finally, task-calibrated emulated $K{=}100$ flips yield $0/20$ real-robot successes, versus $14/20$ clean and $16/20$ global-random. Weight integrity is therefore a security boundary for embodied foundation models. Code is included as ancillary material.
関連論文
- TrustVLA:メカニズムに基づく推論時防御による視覚言語行動モデルのバックドア対策VLA/セキュリティ
- 信頼された想像への攻撃:想像してから行動する世界モデルに対するオラクルレベルの整合性攻撃VLA/セキュリティ
- 軌道レベルでのリダイレクション攻撃:視覚言語行動モデルに対するVLA/セキュリティ