大規模言語モデルにおける二項式の語順選好の行動的・表象的証拠
Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models
大規模言語モデルが言語的な二項式(例:men and women)の語順選好をどの程度捉えているかを、多言語データセットとプロービングを用いて分析し、モデルの内部表現が選好の強さを部分的に符号化していることを示した。
著者: Zhiqing Yang, Yilun Liu, Yunpu Ma, Volker Tresp, Hinrich Schütze
分類: cs.CL, cs.LG
原文アブストラクト
Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains underexplored. We investigate this using linguistic binomials, such as men and women, where both word permutations are grammatically valid but exhibit distinct, cross-linguistic variations in conventionality. We formalize binomial ordering as a distributional alignment problem, and construct a multilingual dataset of 600 binomial pairs across 8 languages. With categorical and distributional metrics, we measure and compare the corpus-derived preferences with model-induced ordering probabilities of 6 open-weight LLMs. While models often behaviorally recover the dominant corpus-preferred order, particularly for strongly conventionalized pairs, they align less well with the exact corpus preference distributions. This suggests that apparent directional order overstates how faithfully LLMs capture the statistical nuances of language use. Sparse probing verifies that the concept of preference strength is partially encoded among middle-to-late layers, and steering along probe-derived directions alters model-induced ordering distributions, demonstrating that the statistical behavioral preference of LLMs can be mechanistically measured and manipulated via internal representations.
関連論文
- 言語モデルに対するダッチブック言語モデル
- ループ型言語モデルにおける再帰計算の割り当て言語モデル
- 言語モデルの潜在軌跡における不変推論方向言語モデル