# 2026-06-03 - Cross-run leaderboard update for the published `step_7000` release This note records the later five-way cross-run re-evaluation that was completed after the original `2026-06-01` release. ## What changed The repo-native CPU benchmark suite `configs/eval/20260521_pretrain_minimal_en_it_webwiki_step11000.yaml` was rerun across five previously selected winner checkpoints from distinct GPT-2-small EN/IT web/wiki run families. Primary decision metric: - `val_loss_mixed` lower is better Final leaderboard: 1. `earlydecay7000 step_7000` - `val_loss_mixed=5.2158` 2. `earlydecay3500 step_7000` - `val_loss_mixed=5.2358` 3. `lr2e4 cosine step_7000` - `val_loss_mixed=5.3558` 4. `lr2e4 WSD step_11000` - `val_loss_mixed=5.3576` 5. `lr1e4 cosine step_14000` - `val_loss_mixed=5.4493` ## Reading The winner stayed unchanged. That means the published `earlydecay7000 step_7000` release remains the current best benchmark checkpoint across the comparable public GPT-2-small EN/IT web/wiki slice tracked from this workspace. This should still be read narrowly: - it is the best current checkpoint under the comparable benchmark protocol - it is not a claim that free-form generations are already strong - qualitative outputs remain weak and repetitive across the leaderboard family ## Added supporting artifacts - `cross_run_leaderboard_report_20260603.md` - `cross_run_leaderboard_comparison_20260603.json` - `cross_run_leaderboard_summary_20260603.json`