Instructions to use Shamima/babylm-2026-multilingual-uniform-100M-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Shamima/babylm-2026-multilingual-uniform-100M-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Shamima/babylm-2026-multilingual-uniform-100M-v2")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Shamima/babylm-2026-multilingual-uniform-100M-v2") model = AutoModelForCausalLM.from_pretrained("Shamima/babylm-2026-multilingual-uniform-100M-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Shamima/babylm-2026-multilingual-uniform-100M-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Shamima/babylm-2026-multilingual-uniform-100M-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Shamima/babylm-2026-multilingual-uniform-100M-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Shamima/babylm-2026-multilingual-uniform-100M-v2
- SGLang
How to use Shamima/babylm-2026-multilingual-uniform-100M-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Shamima/babylm-2026-multilingual-uniform-100M-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Shamima/babylm-2026-multilingual-uniform-100M-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Shamima/babylm-2026-multilingual-uniform-100M-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Shamima/babylm-2026-multilingual-uniform-100M-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Shamima/babylm-2026-multilingual-uniform-100M-v2 with Docker Model Runner:
docker model run hf.co/Shamima/babylm-2026-multilingual-uniform-100M-v2
BabyLM 2026 — MultiLingual track baseline v2 (byte-premium-uniform, WSD)
Iteration v2 of the byte-premium-uniform trilingual baseline. Same
architecture, same tokenizer, same mixture as v1
(Shamima/babylm-2026-multilingual-uniform-100M), but trained with a
Warmup-Stable-Decay schedule rather than cosine, on the same 100M
ref-token budget at the same per-step compute. The motivation: v1's cosine
schedule decayed to 10% of peak by ~90% through training while loss was
still falling; WSD holds at peak and only decays in the last 25%, which lets
more of the budget land at productive learning rates.
Architecture
- Llama (HF
LlamaForCausalLM) — RoPE, RMSNorm, SwiGLU, no biases, tied embeddings - 12 layers · 768 hidden · 12 heads · 2048 FFN
- 1024 sequence length
- 110,119,680 parameters
- Tokenizer: joint byte-level BPE 32 768 (same as v1; reused so the two are directly comparable)
Training
- Data: BabyBabelLM 2026 100M tier (EN/NL/ZH); full corpora loaded in memory and shuffled
- Mixture: byte-premium-uniform via deficit-driven selection (1/3 of reference tokens per language)
- Optimiser: AdamW (β1=0.9, β2=0.95, wd=0.1)
- LR: 6e-4 peak, WSD schedule (warmup 200 → constant peak → linear 25% decay tail to 6e-5)
- Compute: 4× NVIDIA A10G (23 GB), bf16, DDP, micro-batch 16 × grad-accum 2 (eff. batch 128 sequences = 131k tokens/step)
- Tokens consumed at this checkpoint: 100,016,896 byte-premium-adjusted reference tokens (= 1 epoch over the corpus)
- Per-language epochs at this checkpoint: ~1.0 each (well within the BabyLM ≤10-epoch cap)
Revisions
19 fast-eval branches: chck_1M, chck_2M, …, chck_9M, chck_10M, chck_20M, …, chck_90M, chck_100M.
main is chck_100M.
How to evaluate
git clone https://github.com/babylm-org/babylm-eval
cd babylm-eval/multilingual
bash scripts/zeroshot_model.sh --model_name Shamima/babylm-2026-multilingual-uniform-100M-v2
bash scripts/zeroshot_model_fast_all.sh --model_name Shamima/babylm-2026-multilingual-uniform-100M-v2
Comparison vs v1
See https://github.com/silvererudite/bb-lm-challenge-sub for the iteration log, scaffold, and ablation configs.
Citation
@misc{babylm-2026-uniform-v2,
title = {BabyLM 2026 MultiLingual baseline v2 (WSD schedule)},
author = {Hossain, Shamima},
year = {2026},
url = {https://huggingface.co/Shamima/babylm-2026-multilingual-uniform-100M-v2}
}
- Downloads last month
- 1,440