日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
LLMロバスト性arXiv:2610.02432

大規模言語モデルの入力系列変動に対するロバスト性の評価と改善

Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

シェア:XThreadsFacebookLINEはてブBluesky

LLMの敵対的入力変動に対するロバスト性を評価・改善する手法を提案し、プロンプトインジェクションやトロイの木馬、LLM-as-a-Judgeへの攻撃、MCPエージェントの防御までを扱う。

詳しい要約

1. どんなもの?

- LLM の入力系列変動に対する堅牢性を評価・改善する博士論文 - 対象: prompt injection, trojans (backdoors), 自動品質指標の操作 - 提案: 生成堅牢性指標 R_stab(f), 適応的進化型 black-box 攻撃 ASA, 防御 bypass の体系化, 異種モデル committee, MCP 向け AttestMCP と Commit Boundary - 実装: JudgeGuard, TrojanArmor, MCPSec

2. 先行研究と比べてどこがすごい?

- 従来の攻撃成功率 (ASR) 中心の評価に対し、Jensen-Shannon divergence に基づく生成堅牢性指標 R_stab(f) を提案 - 局所攻撃で V(h) <= 1 - R_class(h) の理論保証を与える点が新しい - LLM-as-a-Judge 向けに ASA を開発し ASR 最大 73.8%、open model 間転移 62.6% を達成 - Trojan Detection Challenge 2023 で surrogate trigger の REASR ~0.99、真の trigger recall ~0.17 (baseline ~0.14) - SaTML CTF 2024 で多層防御 bypass を 4 クラスに体系化し ASR を 90% から 15-25% に低減 - 異種モデル committee で Gemma-3-4B の ASR を 47-55 ポイント低減 (7 モデルで 19.3%) - MCP ベース agentic system で AttestMCP と Commit Boundary を提案し MCPBe…

3. 技術・手法の肝は?

- R_stab(f): 小入力摂動下の per-step 出力分布間の Jensen-Shannon divergence に基づく生成堅牢性指標 - 局所攻撃: V(h) <= 1 - R_class(h) を証明。R_class(h) は決定演算子 h が小摂動下で決定を維持する確率 - 非局所攻撃: calibrated empirical model を提案 - ASA: LLM-as-a-Judge 向け適応的進化型 black-box 攻撃 - 異種モデル committee: 5-7 モデルで ASR を低減 - AttestMCP: HMAC 保護パケットで tool call を attest (1 call あたり 0.1 ms 未満) - Commit Boundary: MCP 向け isolation pattern

4. どうやって有効だと検証した?

- LLM-as-a-Judge で ASA の ASR 最大 73.8%、open model 間転移 62.6% を測定 - Trojan Detection Challenge 2023 データ (Pythia-1.4B) で surrogate trigger の REASR ~0.99、真の trigger recall ~0.17 (baseline ~0.14) - SaTML CTF 2024 で多層防御 bypass 4 クラスを体系化し ASR を 90% から 15-25% に低減 - 異種モデル committee 5-7 で Gemma-3-4B の ASR を 47-55 ポイント低減、7 モデルで 19.3% - MCPBench 847 シナリオで AttestMCP と Commit Boundary により平均 ASR を 53.7% から 12.4% に低減

5. 議論はある?

- 局所攻撃に対する理論保証 V(h) <= 1 - R_class(h) を提示 - 非局所攻撃は calibrated empirical model で扱う - 異種モデル committee の有効性を示すが、モデル数や構成の影響は要旨からは不明 - AttestMCP の HMAC 保護は 0.1 ms 未満のオーバーヘッド - 防御 bypass の 4 クラス体系化は SaTML CTF 2024 に基づく - 限界や今後の課題は要旨からは不明

6. 次に読むべき論文は?

- Jensen-Shannon divergence に基づく生成堅牢性指標 R_stab(f) の関連研究 - LLM-as-a-Judge の堅牢性評価 (ASA の比較対象) - Trojan Detection Challenge 2023 (Pythia-1.4B) - SaTML CTF 2024 - Model Context Protocol (MCP) と MCPBench - 異種モデル committee による防御 - 関連ソフトウェア: JudgeGuard, TrojanArmor, MCPSec

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Narek Maloyan

分類: cs.CR, cs.CL, cs.LG

原文アブストラクト

Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.

PR本紙発行元 EmplifAI