日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.05637

視覚言語モデルにおける視覚的グラウンディングの安全性

Visual Grounding Safety in Vision-Language Models

シェア:XThreadsFacebookLINEはてブBluesky

有害な要求が自由記述応答ではなく点やバウンディングボックスのグラウンディングを求める場合、視覚言語モデルが拒否しにくくなることを示し、グラウンディング形式の拒否データで微調整することで安全性を改善する手法を提案した。

詳しい要約

1. どんなもの?

- Vision-Language Models (VLMs) の visual grounding 出力(点や bounding box)における安全性を研究。 - 直接危害、社会的バイアス、状況的安全性の3つのベンチマークを再構成し、15,401の有害リクエストペアを作成。 - ペアは自由テキスト回答(VQA)と grounding(点または bounding box)要求のみが異なる。 - 5つの VLM を評価し、grounding 要求時の拒否率が VQA より31-59ポイント低いことを発見。 - 安全性システムプロンプトではこのギャップを埋められない。 - ファインチューニング手法を提案し、grounding 拒否率を大幅に改善。

2. 先行研究と比べてどこがすごい?

- 従来の VLM 安全性研究は主に自由テキスト出力を対象とし、grounding 出力チャネルの安全性は系統的に分析されていなかった。 - 本研究は初めて grounding 出力に焦点を当て、VQA との安全性ギャップを定量的に示した。 - 既存の安全性アライメントが grounding 出力に十分適用されていないことを明らかにした。 - 提案するファインチューニング手法は、grounding 拒否率を77-95ポイント改善し、VQA 拒否率も向上させ、grounding 能力を維持しつつ過剰拒否を抑制。

3. 技術・手法の肝は?

- 3つの安全性ベンチマーク(VLSU, BBQ-V, Asimov-2.0)を再構成し、有害リクエストのマッチペア(VQA vs grounding)を15,401組作成。 - 5つの VLM で評価し、拒否率の差を測定。 - ファインチューニング手法を提案:grounding 形式の拒否データ、能力 grounding データ、自己蒸留した良性データを組み合わせ。 - これにより grounding 拒否率を向上させ、過剰拒否を抑制。 - 表現分析により、ファインチューニングが有害リクエストを拒否方向に移動させ、特に grounding で顕著であることを示す。

4. どうやって有効だと検証した?

- 5つの VLM で評価し、grounding 拒否率が VQA 拒否率より31-59ポイント低いことを確認。 - 提案手法を Qwen3-VL-8B と VisionReasoner-7B に適用し、VLSU と BBQ-V で grounding 拒否率を77-95ポイント改善。 - ホールドアウトの Asimov-2.0 ドメインで64-85ポイント改善。 - VQA 拒否率も向上し、grounding 能力を維持し、過剰拒否を限定的に抑えることを確認。 - 表現分析により、ファインチューニングが有害リクエストを拒否方向に移動させることを検証。

5. 議論はある?

- 安全性システムプロンプトでは grounding と VQA の拒否率ギャップを埋められないことを指摘。 - ファインチューニングが grounding 拒否率を大幅に改善するが、過剰拒否のバランスや一般化可能性について議論の余地がある。 - 表現分析の結果、ファインチューニングが有害リクエストを拒否方向に移動させ、良性リクエストは無害参照付近に留まることを示す。 - 限界や倫理的影響については要旨からは不明。

6. 次に読むべき論文は?

- VLSU, BBQ-V, Asimov-2.0 の各ベンチマーク論文。 - Qwen3-VL-8B, VisionReasoner-7B のモデル論文。 - 関連手法:visual grounding, safety alignment, fine-tuning, self-distillation。 - 同分野の定番:Vision-Language Models の安全性に関する研究(例:VLGuard, Safety Benchmarks for VLMs)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Erfan Shayegani, Kundan Krishna, Yue Dong, Nael Abu-Ghazaleh, Leon Gatys, Shruti Palaskar

分類: cs.AI, cs.CR, cs.CV, cs.CY, cs.LG

原文アブストラクト

Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We study visual grounding safety by repurposing three safety benchmarks spanning direct harm (VLSU), social bias (BBQ-V), and situational safety (Asimov-2.0) into 15,401 matched pairs of harmful requests that differ only in the requested output: a free-text answer (VQA) or a grounding (point or bounding box). Across five VLMs, models that refuse a harmful request posed as a question often comply when the same request asks for a grounding: averaged over models, grounding refusal trails VQA refusal by 31-59 percentage points, depending on the domain, and safety system prompts do not close this gap. We propose a fine-tuning approach that combines grounding-form refusals with capability grounding data and self-distilled benign data to counter over-refusal. For Qwen3-VL-8B and VisionReasoner-7B, it improves grounding refusal by 77-95 percentage points on VLSU and BBQ-V and by 64-85 points on the held-out Asimov-2.0 domain, while also improving VQA refusal, preserving grounding capability, and keeping over-refusal limited. Representation analysis shows that fine-tuning moves harmful requests toward each model's refusal direction, most strongly for grounding, while leaving benign requests near the harmless reference.

関連論文

PR本紙発行元 EmplifAI