--- license: apache-2.0 tags: [time-series, forecasting, foundation-model, tokenizer-ablation, toto-2-recipe] --- # TSFM tokenizer ablation — Toto-2-4m-recipe clone (v10) VQ-VAE OPTIMIZATION ablation (codebook health + next-code objective vs raw patches) on a ~3.6M-param decoder-only patched transformer trained on the OFFICIAL TempoPFN synthetic prior generators (cloned from automl/TempoPFN; GP/KernelSynth families capped at 6,144 steps, so the long pool uses the cheap families only) with contiguous patch masking, a 9-level quantile head, NorMuon+AdamW, and index RoPE at time-scaled positions. Horizon decoding is fixed patch-32 in every arm; only the CONTEXT tokenizer varies: | arm | tokenizer | history | ctx tokens | |-----|-----------|---------|-----------| | T0 | fixed-32 (control) | 4,096 | 128 | | T1 | pyramid, iso-context | 4,096 | 44 | | T2 | pyramid, iso-token | 16,384 | 128 | | T3 | adaptive equal-surprise | 16,384 | 128 | Each subfolder is one (arm, seed) run: `model.pt` (final), `ckpt_15000.pt` (rank-stability snapshot), `config.json`, `results.json` (dev-GIFT CRPS + long-horizon probe). Dev metrics use a fixed 14-task GIFT-Eval subset — NOT the full leaderboard; treat numbers as ablation-internal, not comparable to published GIFT scores. Generated by the v10 experiment notebook. ## Results ``` === transfer curve: GM-CRPS by checkpoint (dev-14, mean over seeds) === gm_avg2 gm_avg2_std gm@7500 gm@15000 gm@22500 gm@final gm_final_std long_season params_m perp probe_amp n arm N3_vqce_nobonus_nat 0.1812 NaN 0.1979 0.1774 0.1875 0.1749 NaN 0.2454 3.8780 549.6537 0.2580 1 N1_vqce_nat 0.1909 0.0109 0.2121 0.2052 0.1983 0.1835 0.0049 0.2550 3.8780 620.2301 0.2280 3 N0_raw_nat 0.2037 0.0042 0.2011 0.1938 0.2014 0.2061 0.0048 0.2635 3.5802 NaN 0.2253 3 N2_vqce_multi_nat 0.2039 0.0189 0.1913 0.1943 0.2154 0.1925 0.0034 0.2598 4.2743 597.4680 0.2235 2 --- tripwire checks --- N3_vqce_nobonus_nat vs N1_vqce_nat: final gap=0.0086 -> UNRESOLVED (< 2*sigma) N1_vqce_nat vs N0_raw_nat: final gap=0.0226 -> RESOLVED N0_raw_nat vs N2_vqce_multi_nat: final gap=-0.0136 -> UNRESOLVED (< 2*sigma) N0_raw_nat: TRANSFER INVERSION — best at ckpt #2 (0.1938) vs final (0.2061); synthetic overfit persists N2_vqce_multi_nat: TRANSFER INVERSION — best at ckpt #1 (0.1913) vs final (0.1925); synthetic overfit persists ```