日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
解釈可能性/グロッキングarXiv:2609.20166

全列挙可能なTransformer:遅延汎化科学のための計測器

Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization

シェア:XThreadsFacebookLINEはてブBluesky

12Kパラメータの小型Transformerで確立した3つのタスク法則を、12K〜50M(4000倍)で再測定し、天井則と遅延則は保存されるが重み減衰則は変形することを事前登録研究で示した。

詳しい要約

1. どんなもの?

完全に列挙可能なタスク上で訓練された極小の transformer を、grokking(遅延汎化)研究のための「科学計測器」として位置づける論文。 - 全入力の評価、汎化上限の厳密計算、数百 seed の低コスト実行が可能な regime を提案。 - この regime が近似設定では得られない 4 つの能力を持つと主張。 - 厳密で反証可能な汎化上限 - 他を証明可能に固定しつつ 1 つの構造変数を操作する task surgery - 全 weight の直接観測 - 多数 seed の survival-time 統計により「grok しない」を censored observation として扱う - 事前登録した conservation 研究で、12K で確立した 3 つの task-side law を 12K/1M/50M(4,000x スパン、360 runs + 44-run control)で再測定。

2. 先行研究と比べてどこがすごい?

近似設定の grokking 研究と比べ、完全列挙可能 regime は厳密な汎化上限・task surgery・全 weight 観測・survival-time 統計を可能にする点が異なる。 - 「10^4 パラメータで得た法則はそれ以上で意味を持たない」という反論に対し、事前登録の conservation 研究で応答。 - 12K で確立した recoverability-ceiling law と role-conflict delay law は 12K/1M/50M で保存(ceiling 違反 0/144 Holm-corrected、各スケールで Spearman rho >= 0.75、permutation p < 1e-4)。 - weight-decay response law はスケールとともに系統的に変形(steepening)。 - 50M の role-conflict deficit は learning-rate 調整と 3 倍 budget でも残存(事前登録 control)。 - 基準はデータ収集前に凍結され、1 つの law の変形により…

3. 技術・手法の肝は?

完全列挙可能な tiny transformer を model organism として用いる。 - 全入力評価により厳密な汎化上限を計算。 - task surgery で 1 つの構造変数のみを操作し他を証明可能に固定。 - 全 weight を直接観測。 - 多数 seed の survival-time 統計で「grok しない」を censored observation として扱う。 - 事前登録した from-scratch protocol で 12K/1M/50M の 3 スケール(4,000x スパン)を比較。 - 3 つの task-side law(recoverability-ceiling law、role-conflict delay law、weight-decay response law)を再測定。 - Holm-corrected 検定、Spearman rho、permutation p 値、control arm を用いる。

4. どうやって有効だと検証した?

事前登録した conservation 研究で検証。 - 12K で確立した 3 law を 12K/1M/50M で同一 from-scratch protocol により再測定(360 runs + 44-run control arm)。 - ceiling law と delay law は保存:ceiling 違反 0/144(Holm-corrected)、各スケールで Spearman rho >= 0.75、permutation p < 1e-4。 - weight-decay law はスケールとともに系統的に変形(steepening)。 - 50M の role-conflict deficit は learning-rate 調整と 3 倍 budget でも残存(事前登録 control)。 - 基準はデータ収集前に凍結され、1 law の変形によりテストが失敗し得たことを示す。

5. 議論はある?

「10^4 パラメータで特徴づけた法則はそれ以上で意味を持つか」という反論が想定されている。 - 事前登録 conservation 研究により、ceiling law と delay law は 4,000x スパンで保存されるが、weight-decay law は変形することを示す。 - 変形自体も lawful かつ measurable であると主張。 - 完全列挙可能 transformer を delayed generalization の task-side law の model organism として正当化。 - 限界や未解決点の詳細は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として grokking、delayed generalization、transformer の scaling law、weight decay、learning rate scheduling に関する研究が次に読むべき候補。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yoshiyuki Ootani

分類: cs.LG

原文アブストラクト

Tiny transformers trained on fully-enumerable tasks occupy an unusual position in the study of grokking: every input can be evaluated, every generalization ceiling can be computed exactly, and hundreds of seeds cost minutes. We argue this regime is a scientific instrument with four capabilities that approximate settings cannot offer: (a) exact, falsifiable generalization ceilings; (b) task surgery that manipulates one structural variable while provably fixing all others; (c) direct observation of every weight; and (d) survival-time statistics over many seeds that recast "does not grok" as a censored observation. The obvious objection is that laws characterized at 10^4 parameters may not mean anything beyond them. We answer it with a preregistered conservation study: three task-side laws established at 12K parameters -- a recoverability-ceiling law, a role-conflict delay law, and a weight-decay response law -- are re-measured under an identical from-scratch protocol at 12K, 1M, and 50M parameters (a 4,000x span; 360 runs plus a 44-run control arm). The ceiling law and the delay law are conserved (0/144 Holm-corrected ceiling violations; Spearman rho >= 0.75 at every scale, permutation p < 1e-4), while the weight-decay law deforms systematically, steepening with scale. Preregistered controls show the 50M role-conflict deficit survives learning-rate adjustment and a tripled budget. Conservation was tested against criteria frozen before data collection, and one law's deformation shows the test could have failed. These results license the fully-enumerable transformer as a model organism for the task-side laws of delayed generalization: what it measures exactly, larger models largely obey -- and where they deviate, the deviation is itself lawful and measurable.

PR本紙発行元 EmplifAI