# 2026-06-06 - Release note for `step_11000` from the short-fast-decay 11k web/wiki run ## Release candidate - run: `20260603_fresh-gpt2small-lr2e4-bs6-wsd-shortfastdecay11k-final5e6-webwiki` - config: `configs/testing/20260603_fresh-gpt2small-lr2e4-bs6-wsd-shortfastdecay11k-final5e6-webwiki.yaml` - chosen checkpoint: `/mnt/apps/llm-nanochat/checkpoints/20260603_fresh-gpt2small-lr2e4-bs6-wsd-shortfastdecay11k-final5e6-webwiki/step_11000.pt` - intended HF repo: `nazdef/gpt2small-en-it-nanochat-lr2e4-bs6-wsd-shortfastdecay11k-final5e6-webwiki-step11000` ## Why this checkpoint This run finished cleanly at `11000/11000`, and the final checkpoint is both: - the best saved online validation checkpoint: - `validation_loss=3.8305816711` - `validation_perplexity=46.0893392775` - the best checkpoint under the repo-native CPU benchmark across `7000/8000/9000/10000/11000` Primary benchmark ranking by `val_loss_mixed` lower-is-better: 1. `step_11000` - `mixed=5.2270` - `en=4.9931` - `it=4.0279` 2. `step_7000` - `mixed=5.2277` - `en=5.0606` - `it=4.0280` 3. `step_10000` - `mixed=5.2677` 4. `step_8000` - `mixed=5.3010` 5. `step_9000` - `mixed=5.3963` The win over `step_7000` is real but tiny: - `delta mixed (11000 - 7000) = -0.0006882` Supporting signals still lean toward `11000` as the preserve/publish point for this run: - `loop_rate`: `0.40` vs `0.60` at `7000` - `distinct_2`: `0.4990` vs `0.4338` at `7000` - `ppl_en`: `147.39` vs `157.69` at `7000` - `ppl_it`: effectively tied (`56.1452` vs `56.1479`) ## Cross-run reading Against the comparable public champion checkpoint tracked from this workspace: - champion: `20260530_fresh-gpt2small-lr2e4-bs6-wsd-earlydecay7000-final5e6-webwiki step_7000` - champion `val_loss_mixed=5.2158` - this release `val_loss_mixed=5.2270` - delta vs champion: `+0.0112` Operational reading: - this `step_11000` is the best checkpoint of the short-fast-decay run - it is also the most stable-looking checkpoint among the recent short-fast/new candidates in this family - it still does **not** take the overall crown from `earlydecay7000 step_7000`; it misses by a small but real margin on the primary benchmark metric ## Caveats - this is a comparative/operational win, not a claim of polished free-form generation quality - EN improves more clearly than IT - generations remain repetitive enough that "best recent candidate" and "best overall model" are not the same claim