プレプリント / バージョン1

Surface_Key_Conditioned_Successor_Prediction_in_Small_Character_Level_GPT_Models.pdf

##article.authors##

DOI:

https://doi.org/10.51094/jxiv.5581

キーワード:

small language model、 character-level GPT、 shortcut learning、 cross-corpus transfer、 instructional corpus

抄録

An autoregressive model may appear to learn an abstract corpus order when a boundary string alone suffices.
We tested these explanations using three separately worded synthetic corpora. Shared concept labels were analytical
metadata, not common identifiers presented to the model. A 2.67M-parameter character-level GPT was trained
from scratch on six directed-pair-coverage permutations, yielding 90 runs and 270 checkpoints. At iteration 1500,
within-corpus prompts ranked the trained successor first in 360/360 cases at each of k = 32 and k = 64; cross-corpus
prompts at k = 64 produced 210/720 top-1 cases and a mean margin of −0.090. Because a from-scratch character
model may fail to identify a source concept across rewordings, cross-corpus failure is diagnostic rather than primary
evidence. In a 2×2 probe, the top-1 effect of suffix match was 0.621 versus 0.042 for operational content-source match.
With mismatched suffixes, predictions shifted to the suffix donor’s successor in 764/1080 and 750/1080 cases. Crossunit
question substitution selected the donor-cue successor in 1078/1080 cases. Explicit retrieval and variable-order
baselines reproduced the suffix-dependent pattern. Thus the observed effect does not require abstract order learning or
a GPT-specific mechanism, although abstract representations are not ruled out.

利益相反に関する開示

なし

ダウンロード *前日までの集計結果を表示します

ダウンロード実績データは、公開の翌日以降に作成されます。

引用文献

Warstadt, A., Mueller, A., Choshen, L., Wilcox, E., Zhuang, C., Ciro, J., Mosquera, R., Paranjape, B., Williams, A., Linzen, T. and Cotterell, R.: Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora, Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning, Singapore, Association for Computational Linguistics, pp. 1–34 (online), DOI: 10.18653/v1/2023.conll-babylm.1 (2023).

Eldan, R. and Li, Y.: TinyStories: How Small Can Language Models Be and Still Speak Coherent English?, arXiv preprint arXiv:2305.07759 (2023).

Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Behl, H. S., Wang, X., Bubeck, S., Eldan, R., Kalai, A. T., Lee, Y. T. and Li, Y.: Textbooks Are All You Need, arXiv preprint arXiv:2306.11644 (2023).

Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S. and Lee, Y. T.: Textbooks Are All You Need II: phi-1.5 Technical Report, arXiv preprint arXiv:2309.05463 (2023).

Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L. and Van Der Wal, O.: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 202, PMLR, pp. 2397–2430 (online), available from ⟨https://proceedings.mlr.press/v202/biderman23a.html⟩ (2023).

Saphra, N. and Lopez, A.: Understanding Learning Dynamics of Language Models with SVCCA, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, Minnesota, Association for Computational Linguistics, pp. 3257–3267 (online), DOI: 10.18653/v1/N19-1329 (2019).

Tirumala, K., Markosyan, A. H., Zettlemoyer, L. and Aghajanyan, A.: MemorizationWithout Overfitting: Analyzing the Training Dynamics of Large Language Models, Advances in Neural Information Processing Systems, Vol. 35, pp. 38274–38290 (2022).

Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C. and Carlini, N.: Deduplicating Training Data Makes Language Models Better, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, Association for Computational Linguistics, pp. 8424–8445 (online), DOI: 10.18653/v1/2022.acl-long.577 (2022).

Carlini, N., Tram`er, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U´ ., Oprea, A. and Raffel, C.: Extracting Training Data from Large Language Models, 30th USENIX Security Symposium (USENIX Security 21), USENIX Association, pp. 2633–2650 (2021).

Huang, J., Yang, D. and Potts, C.: Demystifying Verbatim Memorization in Large Language Models, arXiv preprint arXiv:2407.17817 (2024).

Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M. and Wichmann, F. A.: Shortcut Learning in Deep Neural Networks, Nature Machine Intelligence, Vol. 2, pp. 665–673 (online), DOI: 10.1038/s42256-020-00257-z (2020).

Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R. and Smith, N. A.: Annotation Artifacts in Natural Language Inference Data, Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), Association for Computational Linguistics, pp. 107–112 (online), DOI: 10.18653/v1/N18-2017 (2018).

McCoy, R. T., Pavlick, E. and Linzen, T.: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, pp. 3428–3448 (online), DOI: 10.18653/v1/P19-1334 (2019).

Kaushik, D., Hovy, E. and Lipton, Z. C.: Learning the Difference That Makes a Difference with Counterfactually-Augmented Data, International Conference on Learning Representations, (online), available from ⟨https://openreview.net/forum?id=Sklgs0NFvr⟩ (2020).

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł. and Polosukhin, I.: Attention Is All You Need, Advances in Neural Information Processing Systems 30, pp. 5998–6008 (online), available from ⟨https://proceedings.neurips.cc/paper/7181-attention-is-all-you-need⟩ (2017).

Geva, M., Schuster, R., Berant, J. and Levy, O.: Transformer Feed-Forward Layers Are Key-Value Memories, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic, Association for Computational Linguistics, pp. 5484–5495 (online), DOI: 10.18653/v1/2021.emnlp-main.446 (2021).

Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S. and Olah, C.: In-context Learning and Induction Heads, arXiv preprint arXiv:2209.11895 (2022).

Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L. and Lewis, M.: Generalization through Memorization: Nearest Neighbor Language Models, International Conference on Learning Representations, (online), available from ⟨https://openreview.net/forum?id=HklBjCEKvH⟩ (2020).

Karpathy, A.: nanoGPT: The Simplest, Fastest Repository for Training/Fine-Tuning Medium-Sized GPTs, Software repository (2022).

ダウンロード

公開済


投稿日時: 2026-07-16 01:40:59 UTC

公開日時: 2026-08-24 09:16:34 UTC
研究分野
情報科学