日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
解釈可能性arXiv:2609.31603

信念自己蒸留によるユーザーモデル抽出

User Model Extraction via Belief Self-Distillation

シェア:XThreadsFacebookLINEはてブBluesky

LLMが内部に持つユーザー信念を読み書き可能な形で抽出する手法を提案し、拒否判断がユーザー意図の推論に依存することを示した。

詳しい要約

1. どんなもの?

- LLMが暗黙的に推論するユーザー属性(belief)を読み書きする枠組み - Belief Self-Distillation (BSD) を提案 - 線形probeとcausal probingを橋渡しする統合read-writeフレームワーク - 凍結LLM自身をteacherとして自然会話からbeliefを蒸留 - 外部アノテーション不要 - 複数モデルファミリでユーザーbeliefを復元し、介入を可能にする

2. 先行研究と比べてどこがすごい?

- 従来のprobingは活性化に存在する情報を抽出するが、因果的役割は直接検証できない - BSDは情報の分離だけでなく、因果的役割を直接テストできる状態を同定 - matched hidden-state steeringよりも実質的に強い介入を実現 - 拒否がリクエストだけでなく推論されたユーザー意図に依存することを発見 - 独立に訓練されたLLMがユーザー表現の共有幾何学に収束するcross-model regularityを発見

3. 技術・手法の肝は?

- 凍結LLMを自身のteacherとしてbeliefを蒸留 - 自然会話から外部アノテーションなしでbeliefを抽出 - コンパクトなユーザー表現を学習し、デコードと書き戻しの両方を可能にする - 線形probeとcausal probingを統合 - 学習された表現をモデルに書き戻すことで因果的介入を実施

4. どうやって有効だと検証した?

- 複数のモデルファミリでBSDがユーザーbeliefを忠実に復元することを確認 - matched hidden-state steeringと比較して介入効果が大幅に強いことを示す - リクエストを固定したままbeliefを変更すると拒否が変化することを検証 - 独立に訓練されたLLM間でユーザー表現の共有幾何学が収束することを発見

5. 議論はある?

- 暗黙のユーザーモデルが読み書き可能な内部状態であることを示す - AI safetyへの直接的含意:モデルが誰と対話していると信じるかに基づいて安全判断を条件付ける - 拒否が推論されたユーザー意図に依存するという発見は安全性設計に影響 - cross-model regularityの意義や限界については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:linear probing, causal probing, hidden-state steering - 関連手法:probing, activation steering, representation engineering - 同分野の定番:LLM interpretability, user modeling, AI safety alignment

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata

分類: cs.LG, cs.CL

原文アブストラクト

Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.

関連論文

PR本紙発行元 EmplifAI