日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
解釈可能性arXiv:2609.32355

言語モデルにおける介入と検出のための勾配と活性化:ターゲット特徴学習

Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models

シェア:XThreadsFacebookLINEはてブBluesky

言語モデル内の特定概念に対応する特徴を、活性化値・活性化勾配・パラメータ勾配と2種の推定器の組み合わせで学習し、検出と因果介入の性能を比較した。

著者: Jonathan Drechsel, Steffen Herbold

分類: cs.CL, cs.LG

原文アブストラクト

Model-internal features can be studied through both their ability to identify a specified concept and their causal effect when manipulated, e.g., through steering or weight editing. A prominent approach to feature learning is Sparse Autoencoders (SAEs), which learn broad feature dictionaries whose relation to particular concepts is typically identified post hoc. However, many interpretability questions are instead hypothesis-driven and concern a concept specified in advance. We study this setting as targeted feature learning, where a single feature is constructed for such a predefined concept. We present a controlled comparison across three model signals (activation values, activation gradients, and parameter gradients) and two estimators (contrastive mean and a learned one-dimensional encoder-decoder), yielding six targeted methods, with CAA and GRADIEND as existing instances and four new methods covering the remaining combinations. We compare these methods against pretrained SAEs across 15 tasks and three language models, evaluating both detection and causal intervention. Across models, the strongest detection performance is achieved by contrastive activation value methods, whereas the strongest intervention performance is achieved by gradient-based methods. Overall, our results show that targeted feature quality depends jointly on the model signal and estimator, with detection and intervention capturing complementary properties.

関連論文

PR本紙発行元 EmplifAI