日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
モデルマージ/活性化ステアリングarXiv:2610.04283

一次近似ステアリング:重み適応を活性化ステアリングに変換する手法

First-Order Steering: Translating Weight Adaptation into Activation Steering

シェア:XThreadsFacebookLINEはてブBluesky

重み更新を活性化空間のステアリングベクトルとして一次近似する枠組みを提案し、複数行動の同時制御を可能にするモデルマージ手法HeRD-Mergingを開発した。

詳しい要約

1. どんなもの?

- 本論文は、Activation Steeringの新しい手法「First-Order Steering」を提案する。 - これは、重み更新行列の一次近似としてActivation Steeringを定式化し、Steering強度のベクトルでパラメータ化する。 - また、一次近似誤差を最小化するモデルマージ手法「HeRD-Merging」を開発し、より正確なSteeringベクトルを生成する。 - これにより、個別および合成された行動の制御が可能になる。

2. 先行研究と比べてどこがすごい?

- 既存のActivation Steering手法では、複数のターゲット行動を同時に適用するためのSteeringベクトルの合成が困難であった。 - 一方、モデルマージの先行研究では、学習された重み適応によって表現されるターゲット行動を高精度で組み合わせられることが示されている。 - 本研究は、重み適応をActivation Steeringベクトルに変換することで、合成可能なSteeringベクトルを生成し、推論時の行動制御を同時に可能にする点が優れている。 - また、HeRD-Mergingは従来のモデルマージベースラインと同等の性能を維持しつつ、より正確な一次Steeringベクトルを生成する。

3. 技術・手法の肝は?

- Activation Steeringを、Steering強度のベクトルでパラメータ化された重み更新行列の一次近似として定式化する。 - 一次Steeringの近似誤差に関する理論的限界を確立する。 - その近似誤差項を最小化する新しいモデルマージ手順「HeRD-Merging」を開発する。 - これにより、個別および合成行動をより正確に制御するSteeringベクトルを生成する。

4. どうやって有効だと検証した?

- 提案手法が個別および合成行動の制御において、既存のActivation Steering手法よりも正確であることを示した。 - HeRD-Mergingが従来のモデルマージベースラインと同等の性能を達成しつつ、より正確な一次Steeringベクトルを生成することを確認した。 - 具体的な評価指標やデータセットは要旨からは不明。

5. 議論はある?

- 一次近似誤差の理論的限界を確立し、その最小化がSteering精度向上に寄与することを議論している。 - モデルマージとActivation Steeringの統合による利点と、合成行動制御の可能性について言及している。 - 限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されている研究:モデルマージの先行研究、既存のActivation Steering手法。 - 関連手法:Activation Steering、モデルマージ、重み適応。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sri Pranav Kunda, Alexander Kurz, Tomas Dominik, Uri Maoz

分類: cs.LG, cs.AI, cs.CL

原文アブストラクト

Activation steering exploits interpretable directions in the residual stream to enable inference-time manipulation of model behavior. Composing steering vectors to apply multiple target behaviors simultaneously is important in various fields-including AI alignment and safety-but remains a challenge for existing activation steering methods. In contrast, prior work in model merging shows that target behaviors represented by learned weight adaptations can be combined with high accuracy. A method that translates weight adaptations into activation steering vectors could therefore extend prior work in model merging to generate composable steering vectors that better enable simultaneous inference-time behavioral control. For this, we introduce First-Order Steering, a formulation of activation steering as a first-order approximation of weight update matrices parameterized by a vector of steering strengths, and establish theoretical bounds on the approximation error of first-order steering. We then develop a novel model merging procedure, HeRD-Merging, which minimizes the first-order approximation error terms to enable higher first-order steering accuracy. Together, our method produces steering vectors that control both individual and composed behaviors more accurately than existing activation steering methods. Furthermore, HeRD-Merging matches the performance of conventional model-merging baselines, while producing weight adaptations that admit more accurate first-order steering vectors.

PR本紙発行元 EmplifAI