--- license: mit tags: - babylm - babylm-2026 - modernbert - masked-language-model - strict-small language: - en library_name: transformers pipeline_tag: fill-mask --- # ModernBERT-Small (Factorized Linear Embeddings) — BabyLM 2026 Strict-Small This model is a `ModernBERT-Small` masked language model (RoPE, GeGLU, alternating local/global attention, 384 hidden size, 16 layers, 6 attention heads) trained from scratch on the **BabyLM 2026 Strict-Small** 10M-word corpus, as part of an ablation study on parameter-efficient token embedding layers for developmentally-plausible pretraining under the BabyLM Challenge's strict-small compute and data budget. ## Embedding design Standard transformer token embedding tables scale as `vocab_size x hidden_size`, which for this model's 30,522-token vocabulary and 384 hidden size would be an 11.7M-parameter dense lookup table -- a large fraction of the model's total parameter budget under the strict-small constraint. This checkpoint instead uses an **ALBERT-style linear factorization**: tokens are first embedded into a smaller 128-dimensional bottleneck space, then linearly projected up to the model's 384-dimensional hidden size, replacing one large `vocab_size x hidden_size` matrix with two much smaller ones (`vocab_size x 128` and `128 x 384`). The output projection (`tie_word_ embeddings=false`) uses its own separate decoder rather than sharing the input embedding weights. This is one variant in a broader comparison of embedding-layer parameterizations (dense, linear and MLP-style factorization, tensor-train decomposition, deterministic Fourier expansion, and compositional/frequency-adaptive variants) evaluated under identical data, tokenizer, optimizer, and training budget, to isolate the effect of the embedding layer's parameterization on downstream BabyLM evaluation performance. ## Training data BabyLM 2026 Strict-Small corpus (~10M words), tokenized with a byte-level BPE tokenizer trained on the same corpus (vocab size 30,522). No external data, synthetic augmentation, or human annotation beyond the corpus as officially released. ## Usage ```python from transformers import AutoModelForMaskedLM, AutoTokenizer model = AutoModelForMaskedLM.from_pretrained( "remg1997/modernbert-small-modernbert-small-factorized-linear-babylm2026", trust_remote_code=True, ) tokenizer = AutoTokenizer.from_pretrained( "remg1997/modernbert-small-modernbert-small-factorized-linear-babylm2026" ) ``` Each `chck_{N}M` branch of this repository corresponds to a BabyLM-Challenge compliance checkpoint (one per `N` million words of training data seen); `main` points at the final, fully-trained checkpoint. ## Evaluation Evaluated with the official [`babylm-eval`](https://github.com/babylm-org/babylm-eval) harness (BLiMP, EWoK, entity tracking, COMPS, Global PIQA, reading-time correlation, GLUE/SuperGLUE fine-tuning, and Age-of-Acquisition word-surprisal correlation) under the strict-small track.