プレプリント / バージョン1

教科横断型人工教材系列でゼロから学習した 小型GPTにおける単元内・単元間遷移の出力形成過程

##article.authors##

DOI:

https://doi.org/10.51094/jxiv.5677

キーワード:

小型言語モデル、 人工教材系列、 学習過程、 データ配列、 自己回帰生成

抄録

本研究では,教材を模した階層的な日本語系列に反復された単元内・単元間遷移が,事前学習なし
小型GPT の出力へ形成される過程を分析した.国語・算数・生活の3 教科と4 テーマからなる12 単元
を作成し,教科ブロック,テーマ交互,有効単元シャッフル,単元内役割ブロック入替の4 条件を比較し
た.4 条件について,2 モデル規模,3 種類のlayout seed,5 種類のmodel seed を組み合わせた120 runs
を,7 checkpoints で評価した.自由生成とteacher-forced 候補評価を併用し,統計単位はmodel run とし
た.競合系列と次単元が重複する地点を除いた構造選択性は,iteration 700 で4 条件・2 モデル規模の全
120 runs において正となった.教科・テーマ属性に基づく構造化系列の有効シャッフル系列に対する優位
性はHolm 補正後には支持されなかった.単元内評価では,厳密なcompletion-prefix 一致と最初に検出さ
れた役割class を用いた.役割ブロック入替条件における訓練役割への局所選択性は0.892(0.8M)および
0.917(2.67M)であった.また,同一更新回数・同一投入文字token 数に対して,2.67M モデルでは構造
適合がより早いcheckpoint で観測された.以上から,本実験範囲の小型GPT の出力は,コーパス中に反
復された単元内・単元間境界遷移に対応して形成され,属性に基づく構造化系列が有効シャッフル系列よ
り強く形成される傾向は確認されなかった.

利益相反に関する開示

なし

ダウンロード *前日までの集計結果を表示します

ダウンロード実績データは、公開の翌日以降に作成されます。

引用文献

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. and Polosukhin, I.: Attention Is All You Need, Advances in Neural Information Processing Systems, Vol. 30 (2017).

Li, K., Hopkins, A. K., Bau, D., Viegas, F. B., Pfister, H. and Wattenberg, M.: Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task, The Eleventh International Conference on Learning Representations (2023).

Power, A., Burda, Y., Edwards, H., Babuschkin, I. and Misra, V.: Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets, arXiv preprint arXiv:2201.02177 (2022).

Murty, S., Sharma, P., Andreas, J. and Manning, C. D.: Grokking of Hierarchical Structure in Vanilla Transformers, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Association for Computational Linguistics, pp. 439–448 (online), DOI: 10.18653/v1/2023.aclshort.38 (2023).

Tian, Y., Wang, Y., Chen, B. and Du, S. S.: Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer, Advances in Neural Information Processing Systems, Vol. 36, pp. 71911–71947 (2023).

Qian, P., Naseem, T., Levy, R. and Fernandez Astudillo, R.: Structural Guidance for Transformer Language Models, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, pp. 3735–3745 (2021).

Papadimitriou, I. and Jurafsky, D.: Injecting Structural Hints: Using Language Models to Study Inductive Biases in Language Learning, Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 8402–8413 (2023).

Ahuja, K., Balachandran, V., Panwar, M., He, T., Smith, N. A., Goyal, N. and Tsvetkov, Y.: Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers, Transactions of the Association for Computational Linguistics, Vol. 13, pp. 121–141 (2025).

Kiyomaru, H., Oda, Y., Kodama, T., Liu, C. and Kawahara, D.: Scaling Data-Constrained Language Models with Synthetic Data, Findings of the Association for Computational Linguistics: EACL 2026, Association for Computational Linguistics, pp. 1002–1016 (online), DOI: 10.18653/v1/2026.findings-eacl.52 (2026).

Bengio, Y., Louradour, J., Collobert, R. and Weston, J.: Curriculum Learning, Proceedings of the 26th Annual International Conference on Machine Learning, pp. 41–48 (2009).

Chan, S. C. Y., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A., Richemond, P. H., McClelland, J. L. and Hill, F.: Data Distributional Properties Drive Emergent In-Context Learning in Transformers, Advances in Neural Information Processing Systems, Vol. 35, pp. 18878–18891 (2022).

Singh, A., Chan, S., Moskovitz, T., Grant, E., Saxe, A. and Hill, F.: The Transient Nature of Emergent In-Context Learning in Transformers, Advances in Neural Information Processing Systems, Vol. 36 (2023).

Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. and Amodei, D.: Scaling Laws for Neural Language Models, arXiv preprint arXiv:2001.08361 (2020).

Bengio, S., Vinyals, O., Jaitly, N. and Shazeer, N.: Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks, Advances in Neural Information Processing Systems, Vol. 28 (2015).

Ranzato, M., Chopra, S., Auli, M. and Zaremba, W.: Sequence Level Training with Recurrent Neural Networks, International Conference on Learning Representations (2016).

Karpathy, A.: nanoGPT, GitHub repository (2022).

ダウンロード

公開済


投稿日時: 2026-07-21 07:25:51 UTC

公開日時: 2026-08-07 10:45:04 UTC
研究分野
情報科学