日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
LLM安全性arXiv:2610.07125

埋め込み摂動によるオープンウェイトLLMのジェイルブレイク

Jailbreaking Open-Weight LLMs via Random Embedding Perturbations

シェア:XThreadsFacebookLINEはてブBluesky

プロンプトの埋め込みベクトルにガウスノイズを加えるだけで、6つのオープンウェイトLLMを高速かつ低コストでジェイルブレイクできる手法PEVを提案した。

詳しい要約

1. どんなもの?

- オープンウェイトLLMの安全性脆弱性を暴露する研究 - 提案手法PEVは、プロンプトの埋め込みベクトルに独立なガウスノイズを加えるだけでjailbreakを実現 - 勾配計算やプロンプトごとの最適化、モデル内部重みの変更が不要 - 6つの一般的なオープンウェイトLLM(サイズ多様)でJailbreakBenchベンチマークを用いて評価 - 有害または安全でない応答を一貫して引き出すことに成功

2. 先行研究と比べてどこがすごい?

- 従来のjailbreak手法は勾配計算、プロンプトごとの最適化、内部重みの変更を必要とし、計算コストが高い - PEVは埋め込みベクトルにガウスノイズを加えるだけのシンプルで高速な手法 - 最初の成功攻撃までの平均計算コストが従来手法より最大1桁少ない - 新規プロンプトに対する最初のjailbreakが全テストモデルで通常1分以内に達成 - JailbreakBenchの全プロンプト・全モデルで安全でない応答を生成。他のテスト手法は同等の結果を得られず、実行時間も長い

3. 技術・手法の肝は?

- プロンプトの埋め込みベクトル表現に独立なガウスノイズを加える - 追加の操作は不要 - 安全でない応答を生成するために、この分布から加法ノイズを繰り返しサンプリング - 勾配計算、プロンプトごとの最適化、モデル内部重みの変更を必要としない - シンプルで高速、低コストなjailbreak手法

4. どうやって有効だと検証した?

- 6つの一般的なオープンウェイトLLM(サイズ多様)を対象に実験 - JailbreakBenchベンチマークデータセットを使用 - 最初の成功攻撃までの平均計算コストを測定し、従来手法と比較 - 新規プロンプトに対する最初のjailbreakまでの時間を計測(全モデルで通常1分以内) - 全モデル・全プロンプトで安全でない応答を生成することを確認 - 他のテスト手法との比較でPEVの優位性を検証

5. 議論はある?

- 埋め込みベクトルの摂動下でのLLMの振る舞いの理解が重要な研究方向であると主張 - 摂動は主要なセキュリティリスクである一方、モデルの動的挙動を探る有用なツールにもなり得る - 具体的な議論や限界については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:JailbreakBenchベンチマーク、従来のjailbreak手法(勾配計算、プロンプトごとの最適化、内部重み変更を用いるもの) - 関連手法:埋め込みベクトルへの摂動、ガウスノイズを用いた攻撃 - 同分野の定番:LLMの安全性評価、jailbreak攻撃、敵対的摂動

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Abhinav Sudhakar Dubey, Scott Sirri, Vaggos Chatziafratis, C. Seshadhri

分類: cs.CR, cs.AI, cs.CL, cs.LG

原文アブストラクト

While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabilities across six common open-weight LLMs of various sizes that consistently lead to harmful or unsafe responses on the JailbreakBench benchmark dataset. Our proposed attack, Perturbed Embedding Vector (PEV), is a simple and fast "jailbreaking" technique that is cheaper than prior approaches, which typically require gradient computations, per-prompt optimizations, or altering internal weights of the models. PEV just adds independent Gaussian noise in the embedding vector representations of the prompt, with no need for further manipulations. To generate unsafe responses, we repeatedly sample additive noise from this distribution. In our experiments, we observe that the average compute cost to get the first successful attack is up to an order of magnitude less than previous attacks. The first successful jailbreak on a new prompt typically arrives within one minute on every tested model, and PEV generates unsafe responses across all models for all prompts in JailbreakBench. No other tested method achieves such results, despite them taking longer to run. More broadly, we believe that understanding the behavior of LLMs under perturbations in the embedding vectors is an important research direction: while perturbations constitute a major security risk, they can also serve as a valuable tool for exploring the dynamical behavior of such models.

関連論文

PR本紙発行元 EmplifAI