TSFM tokenizer ablation โ Toto-2-4m-recipe clone (v10)
VQ-VAE OPTIMIZATION ablation (codebook health + next-code objective vs raw patches) on a ~3.6M-param decoder-only patched transformer trained on the OFFICIAL TempoPFN synthetic prior generators (cloned from automl/TempoPFN; GP/KernelSynth families capped at 6,144 steps, so the long pool uses the cheap families only) with contiguous patch masking, a 9-level quantile head, NorMuon+AdamW, and index RoPE at time-scaled positions. Horizon decoding is fixed patch-32 in every arm; only the CONTEXT tokenizer varies:
| arm | tokenizer | history | ctx tokens |
|---|---|---|---|
| T0 | fixed-32 (control) | 4,096 | 128 |
| T1 | pyramid, iso-context | 4,096 | 44 |
| T2 | pyramid, iso-token | 16,384 | 128 |
| T3 | adaptive equal-surprise | 16,384 | 128 |
Each subfolder is one (arm, seed) run: model.pt (final), ckpt_15000.pt
(rank-stability snapshot), config.json, results.json (dev-GIFT CRPS +
long-horizon probe). Dev metrics use a fixed 14-task GIFT-Eval subset โ NOT
the full leaderboard; treat numbers as ablation-internal, not comparable to
published GIFT scores. Generated by the v10 experiment notebook.
Results
=== transfer curve: GM-CRPS by checkpoint (dev-14, mean over seeds) ===
gm@7500 gm@15000 gm@22500 gm@final gm_final_std long_season params_m perp probe_amp n
arm
M2_vqce_tpfn 0.2143 0.2073 0.2272 0.2006 0.0048 0.2611 3.9762 305.9201 0.196 2
M0_raw_tpfn 0.2111 0.2178 0.2066 0.2060 0.0044 0.2598 3.5802 NaN 0.208 2
M1_vqtuned_tpfn 0.2371 0.2264 0.2271 0.2140 0.0083 0.2608 3.7130 367.0948 0.320 2
--- tripwire checks ---
M2_vqce_tpfn vs M0_raw_tpfn: final gap=0.0055 -> UNRESOLVED (< 2*sigma)
M0_raw_tpfn vs M1_vqtuned_tpfn: final gap=0.0080 -> UNRESOLVED (< 2*sigma)