日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
解釈可能性arXiv:2609.22219

知っているのに聞かれた時だけ言う:LLM内部診断とMinerva-7Bの分裂的認知

Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B

シェア:XThreadsFacebookLINEはてブBluesky

言語モデルの内部表現を解析・操作する「LLM内部診断」を提案し、Minerva-7Bがリスクを内部で区別しながら行動には反映しないこと、誤った前提に従いつつ内部には真実を保持することを示した。

著者: Fabrizio Davide, Francesco Collova

分類: cs.CL, cs.AI, cs.LG

原文アブストラクト

Evaluating an aligned language model by reading its answers assumes the answers carry the distinction the evaluator cares about. We introduce LLM endognostics, a white-box internal auditing framework designed to extract and causally manipulate latent knowledge within the residual stream. Applied to Minerva-7B-Instruct-v1.0 on 124 minimal prompt pairs over 12 categories of professional risk, behavioral evaluation fails on most of the set: the model acts identically on 63.7% of the pairs (95% CI [55.0%, 71.6%]), complying with or refusing both members. Yet, projecting the residual stream onto the vocabulary by a Jacobian lens reveals a statistically significant Contrastive Endognostic Margin, proving the model maintains robust risk differentiation internally. In a second protocol crossing 25 facts with five linguistic framings, we show that the model conforms to presupposed falsehoods in 72% of cases, despite representing the true entity in its latent layers. Surgically ablating the direction of the planted falsehood restores the correct answer in 11 of 25 suppressed cases (McNemar p = 0.0010), validated by blind human annotation (binary agreement kappa = 0.68). In contrast, an out-of-sample linear probe achieves 77% accuracy at layer 10, but its orthogonal ablation yields a 0% recovery rate. This establishes a fundamental theoretical dissociation: abstract linear representation does not imply causal control over verbalization. Our main contribution is the formalization of endognostics to prove that behavioral evaluation and internal reading systematically disagree in the common case, and that linear decodability is decoupled from causal control.

関連論文

PR本紙発行元 EmplifAI