AuAu: 大規模言語モデルにおける権威主義的整合性の監査のためのベンチマーク
AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models
大規模言語モデル(LLM)の応答における権威主義的傾向を評価する包括的なベンチマーク「AuAu」を提案し、17モデルを評価して権威主義的応答率やシステムプロンプトによる操作可能性を明らかにした。
著者: Andreas Einwiller, Max Klabunde, Florian Lemmerich
分類: cs.CL, cs.AI, cs.LG
原文アブストラクト
The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of whether specific models exhibit or promote authoritarian attitudes. We introduce AuAu, a comprehensive benchmark for assessing the risk of authoritarian tendencies in LLM responses. AuAu combines three evaluation approaches: (i) psychometric questions from 15 human-validated instruments, (ii) vignettes probing intended behavior in concrete situations, and (iii) responses to realistic user prompts. Unlike prior work, AuAu measures not only overall authoritarian alignment but also its established sub-concepts: Authoritarian Aggression, Authoritarian Submission, and Conventionalism. Evaluating 17 models from China, the EU, Russia, and the USA, we find substantial authoritarian response rates on psychometric instruments across all models, though rates drop significantly on more realistic downstream tasks. Moreover, a simple authoritarian system prompt manipulates 15 of 17 models into promoting increased authoritarianism. Our results underscore the need for continued, systematic auditing of LLM-based AI systems to detect and mitigate authoritarian tendencies in their outputs.