日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
解釈可能性arXiv:2610.06903

Transformerの拒否機構におけるコンポーネントと次元のスパース性

Component and Dimension Sparsity in Transformer Refusal Mechanisms

シェア:XThreadsFacebookLINEはてブBluesky

4つのオープンウェイトモデルで拒否ステアリングを分解し、拒否行動が少数の注意・MLPコンポーネントと残差ストリーム次元に集中して符号化されていることを示した研究。

著者: Vincent Siu, Glenn Grant-Richards, Vlad Pavlovich, Yizhou Sun, Dawn Song, Chenguang Wang

分類: cs.CL, cs.AI, cs.LG

原文アブストラクト

Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28--48\% of upstream components, retaining 88--101\% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50\% of residual stream dimensions, retaining 85--98\% of the component-mechanism baseline, consistent with a privileged basis structure. Sparsity thus operates at two levels: which components are steered, and which dimensions within those components carry the signal. Together these findings show that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, providing a foundation for mechanistic understanding of how refusal behaviors are represented and steered. To facilitate reproducibility, we release all code and raw experimental results in https://github.com/wang-research-lab/Refusal_Mechanisms.

関連論文

PR本紙発行元 EmplifAI