Localize-and-Detect: 命令チューニングモデルにおけるタスクレベルポイズニングの監査
Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models
ベースモデルとファインチューニング済みモデルの出力のみを用いて、タスクレベルポイズニングの標的タスクを特定し、偏った内容を検出する2段階のブラックボックス監査手法を提案。
著者: Luze Sun, Cristina Nita-Rotaru, Alina Oprea
分類: cs.CR, cs.LG
原文アブストラクト
Instruction fine-tuning adapts a pretrained language model to follow instructions by training it on instruction--response pairs from a collection of tasks, such as summarization and question answering. Task-level poisoning exploits this task structure to manipulate the fine-tuned model into producing attacker-specified biased content on a particular target task, without requiring an explicit input trigger. Detecting such attacks is challenging because there is no explicit trigger to identify, the target task and biased content are unknown, and benign fine-tuning itself changes model behavior. We introduce Localize-and-Detect, a two-stage black-box auditing method for task-level poisoning that requires only outputs from both the base and fine-tuned models. In the first stage, we localize the target task by identifying candidate tasks on which the fine-tuned and base models have the largest differences in their next-token distributions. In the second stage, we search the shortlisted tasks for biased content that repeatedly appears in the fine-tuned model's responses but not in those of the base model. We evaluate Localize-and-Detect on 216 poisoned models across two model families, varying the target task, poisoning mode, poison budget, and type of biased content. Our evaluation demonstrates that Localize-and-Detect can effectively localize target tasks and detect biased content across a range of poisoning settings and models, with detection that tracks attack success and few false positives on clean models.