日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
Webエージェント/敵対的学習arXiv:2610.08773

AdvSim2Real:適応的プロンプトインジェクションに対するWebエージェントの学習

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

シェア:XThreadsFacebookLINEはてブBluesky

凍結されたWeb世界モデル内でタスクカリキュラム・インジェクション攻撃者・エージェントを共進化させ、未知の攻撃者に対しても堅牢なWebエージェントを訓練する手法を提案。

詳しい要約

1. どんなもの?

- Web agent が第三者ページの injection によりユーザ目標から逸脱する問題に対処。 - AdvSim2Real を提案。 - 凍結された web world model 内で task curriculum、injection adversary、agent を共進化。 - 4B agent の能力と堅牢性を向上。

2. 先行研究と比べてどこがすごい?

- 既存防御は訓練前に固定された injection で fine-tune するため、適応攻撃者に bypass される。 - adversarial training は攻撃者を適応させるがタスクが固定で、解けると学習が止まる。 - AdvSim2Real はタスクと攻撃者を共進化させ、未見の攻撃者にも耐性。

3. 技術・手法の肝は?

- 凍結 web world model 内で共進化。 - curriculum は agent が約半分成功するタスクに報酬。 - adversary は success flip(成功を失敗に変える injection)のみに報酬。 - これにより agent は能力と堅牢性を同時に獲得。

4. どうやって有効だと検証した?

- 4B agent で検証。 - 攻撃あり/なしで completion が上昇。 - 訓練していない frontier-model adversary に対しても堅牢性を維持。 - 能力向上は実ブラウザに転移。 - 150 web tasks で未見 adversary 下の completion が base agent 比 33.6% 相対改善。

5. 議論はある?

- 要旨からは不明。 - 限界や倫理的懸念についての議論は記載なし。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 固定 injection で fine-tune する防御、adversarial training。 - 関連手法: web world model、task curriculum、injection adversary。 - 同分野の定番: WebArena、Mind2Web など web agent ベンチマーク。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth Vepakomma, Nils Lukas

分類: cs.CL, cs.AI, cs.LG

原文アブストラクト

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.

PR本紙発行元 EmplifAI