言語モデルの出力分布サンプリングによる誤差と内部状態の幾何学的関係
A geometric relation of the error introduced by sampling a language model's output distribution to its internal state
GPT型言語モデルが多様なトークンに確率が分散する生成点で単一トークンの変化に敏感であることを幾何学的性質として捉え、トークン埋め込みの幾何のみに依存する1形式を導出。その曲率がチェス推論タスクで意味的に世界モデルと結びつくことを示した。
分類: cs.LG
原文アブストラクト
GPT-style language models are sensitive to single-token changes at generation points where the predicted probability distribution is spread across multiple tokens. Viewing this sensitivity as a geometric property, we derive an $\mathfrak{so}(n)$-valued 1-form that depends only on the geometry of the token embeddings. Despite this purely geometric origin, we show that its curvature is semantically meaningful: On chess reasoning tasks, the curvature couples to the world model of an off-the-shelf instruction-tuned model, with transformations clustering by board region and respecting piece importance. Our findings suggest that token space geometry directly reflects how models internally represent problems.