GDN Paper 120M hdim768 Headwise Q/K A100-kernel TorchAdamW 20k r1
This repository contains a PyTorch Lightning checkpoint for the 120M
MLP-Mixer/GDN language-model experiment m06d23y26-gdn-paper-120m-hdim768-headwiseqk-a100kernel-torchadamw-20k-r1.
Files
last.ckpt: final checkpoint, copied fromstep_20000.ckptin the Perlmutter checkpoint directory.train_config.yaml: training config used for this run.eval_metrics.json: training/eval provenance and scalar metrics.
Provenance
- Source checkout:
/pscratch/sd/j/jwl50/online-mlps/mlp-mixer, commit8e24034. - Training Slurm job:
55019482. - W&B: m06d23y26-gdn-paper-120m-hdim768-headwiseqk-a100kernel-torchadamw-20k-r1.
- Perlmutter checkpoint dir:
/pscratch/sd/j/jwl50/online-mlps/artifacts/checkpoints/mlp-mixer-pile/m06d23y26-gdn-paper-120m-hdim768-headwiseqk-a100kernel-torchadamw-20k-r1.
Metrics
- W&B val PPL:
10.326167120432732. - W&B test PPL:
10.372360659660028. - Pile validation 10% PPL at L=2048:
10.2171724795. - Rare-AR overall PPL at L=2048:
10.2177474204. - Rare-AR selected accuracy:
0.713182175439.
Usage Note
This is not a Transformers-format model. It is a Lightning checkpoint
intended for the mlp-mixer training/eval stack. The existing downloader
convention expects the checkpoint filename last.ckpt.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support