日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
プライバシー保護/敵対的例arXiv:2607.10329v1

視覚言語モデルに対する知覚不能かつ可逆な敵対的例によるプライバシー保護

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語モデルへのテキストベースのプライバシー攻撃から画像を守るため、拡散モデルと可逆ネットワークを組み合わせた高画質で可逆な敵対的例生成フレームワークCloakDiffを提案した。

著者: Qi Lu, Ziqi Zhou, Yufei Song, Zijing Li, Lulu Xue, Minghui Li, Shengshan Hu, Leo Yu Zhang

分類: cs.CV

原文アブストラクト

Vision Language Models (VLMs) offer powerful multimodal ability but also expose users to text-based privacy attacks where adversaries crawl online photos and query VLMs to extract sensitive attributes. Existing reversible adversarial example (RAE) methods protect images in purely visual tasks but fail in multimodal settings, and current adversarial examples on VLMs rely on high frequency noise that severely degrades visual quality. We propose CloakDiff, the first framework for reversible, high fidelity privacy protection against text-based query attacks in VLMs. CloakDiff produces imperceptible adversarial examples by combining diffusion based adversarial editing with an invertible network that embeds the original image for lossless recovery. It perturbs both pixel space embeddings and manipulates latent cross attention maps to ensure strong cross-model and cross-prompt transferability while preserving global visual structure. To further enhance fidelity, we design EDM Heuristic Sampling, a principled diffusion schedule for adversarial guidance. Experiments on multiple datasets and VLMs demonstrate that CloakDiff delivers multimodal privacy preservation with high visual quality and reversibility.