GDN Paper 120M hdim768 Headwise Q/K A100-kernel TorchAdamW 20k r1

This repository contains a PyTorch Lightning checkpoint for the 120M MLP-Mixer/GDN language-model experiment m06d23y26-gdn-paper-120m-hdim768-headwiseqk-a100kernel-torchadamw-20k-r1.

Files

  • last.ckpt: final checkpoint, copied from step_20000.ckpt in the Perlmutter checkpoint directory.
  • train_config.yaml: training config used for this run.
  • eval_metrics.json: training/eval provenance and scalar metrics.

Provenance

Metrics

  • W&B val PPL: 10.326167120432732.
  • W&B test PPL: 10.372360659660028.
  • Pile validation 10% PPL at L=2048: 10.2171724795.
  • Rare-AR overall PPL at L=2048: 10.2177474204.
  • Rare-AR selected accuracy: 0.713182175439.

Usage Note

This is not a Transformers-format model. It is a Lightning checkpoint intended for the mlp-mixer training/eval stack. The existing downloader convention expects the checkpoint filename last.ckpt.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support