日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
敵対的機械学習arXiv:2609.31643

ゲーミング攻撃と学習攻撃に対する情報設計

Information Design Against Gaming and Learning Adversaries

シェア:XThreadsFacebookLINEはてブBluesky

棄権オプション付き二値分類器において、境界付近での棄権と固定率での棄権がBlackwell非比較であり、敵の種類に応じて境界再構成に必要なクエリ数が大きく異なることを理論と実験で示した。

著者: Madhava Gaikwad

分類: cs.LG, cs.AI, cs.CR, cs.GT

原文アブストラクト

A principal who deploys a binary classifier with an abstention option must decide which queries the mechanism abstains on. The right choice depends on the adversary. A gaming adversary already knows the classifier and tries to manipulate features across the boundary, so the principal does best by abstaining on queries close to that boundary. The same boundary-localizing rule is the worst possible choice against a learning adversary who does not know the classifier: each abstention now tells the adversary that the boundary is nearby, which is enough to drive a binary search. We analyze this tension. The two natural defenses, abstaining at a fixed rate and abstaining near the boundary, are Blackwell-incomparable: neither can be simulated by post-processing the other's responses. The number of queries needed to reconstruct the boundary to error $\eps$ is $\tildeΘ(d/\eps)$ under the first defense and $Θ(d \log(1/\eps))$ under the second, where $d$ is the VC dimension of the classifier family and $\tildeΘ$ suppresses factors polylogarithmic in $d$ and $1/\eps$. The first rate is a worst case over query distributions; no reconstruction algorithm can close the gap at the distributions that attain it. We characterize the Pareto frontier between the two defense objectives, and confirm both rates on seven binary-classification tasks spanning tabular, image, and language-model-feature inputs: label-plus-counterfactual access extracts the boundary with up to $200\times$ fewer queries than a published label-only baseline.

PR本紙発行元 EmplifAI