日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.31558

CLIPの埋め込み空間バックドアに対する領域レベルのブラックボックス防御

Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP

シェア:XThreadsFacebookLINEはてブBluesky

CLIPエンコーダの埋め込み空間バックドアに対し、セグメント単位の埋め込み摂動を測って疑わしい領域だけを意味的インペインティングで浄化する、完全ブラックボックスの軽量防御CLIPGuardを提案した。

詳しい要約

1. どんなもの?

- CLIPのembedding-space backdoor攻撃に対する防御手法CLIPGuardを提案。 - 完全black-box設定で動作し、モデルパラメータや勾配、logits、clean validation dataを必要としない。 - 画像をsegmentに分割し、segment-wise embedding perturbationを測定して悪意ある領域を特定。 - 疑わしいsegmentのみをsemantic inpaintingで選択的に浄化し、良性の視覚内容とalignment品質を保持。 - STL-10, ImageNet, 多様なtrigger family (BadCLIP, BadNets, blended, patch-based, typographic) で評価。 - 攻撃成功率を最大1.05%まで低減し、clean accuracyを最大86.34%維持。

2. 先行研究と比べてどこがすごい?

- 既存防御はモデルパラメータ、勾配、logits、clean validation dataへのアクセスを前提とし、現実のblack-box展開では成立しにくい。 - 既存black-box手法は小さいまたはout-of-distributionなtriggerの正確な位置特定が困難。 - CLIPGuardは完全black-boxで動作し、segment単位の摂動測定と選択的浄化により、これらの制約を克服。 - 実験でCleanCLIPやCleanerCLIPなどの既存black-box防御を一貫して上回る性能を示す。

3. 技術・手法の肝は?

- 画像をsegmentに分割し、各segmentのembedding perturbationを測定。 - 摂動が大きいsegmentを悪意ある領域として特定。 - 疑わしいsegmentのみをsemantic inpaintingで浄化し、良性部分はそのまま保持。 - これによりalignment品質を損なわずにbackdoorを無効化。 - 完全black-boxで動作し、モデル内部へのアクセスを必要としない。

4. どうやって有効だと検証した?

- STL-10, ImageNetデータセットで評価。 - BadCLIP, BadNets, blended, patch-based, typographic attacksを含む多様なtrigger familyで検証。 - 攻撃成功率を最大1.05%まで低減。 - clean accuracyを最大86.34%維持。 - CleanCLIPやCleanerCLIPなどの既存black-box防御と比較し、一貫して優位性を確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- CleanCLIP - CleanerCLIP - BadCLIP - BadNets - blended attack - patch-based attack - typographic attack

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ahmed Abdelnaby, Mohamed Elmahallawy

分類: cs.CV, cs.CR

原文アブストラクト

Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git

関連論文

PR本紙発行元 EmplifAI