File size: 312 Bytes
6312baa | 1 2 3 4 5 6 7 8 9 10 11 12 | ---
datasets:
- phonemetransformers/IPA-BabyLM
language:
- en
base_model:
- openai-community/gpt2
---
GPT2 trained on the BabyLM 2024 training set using a BPE tokenizer.
Model trained for [From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes](https://arxiv.org/abs/2410.22906). |