KLD Benchmarks + why Q8_K_XL vs MXFP4 naming

#11
by danielhanchen - opened

Hey folks - for folks asking:

  1. Why did we name the MXFP4 quant "Q8_K_XL" and not MXFP4
  2. Why other folks' "lossless" quants are 155GB vs our Q8_K_XL at 162GB
  3. Is FP8 and Q8_0 equivalent?

FP8 and Q8_0 are NOT bitwise equivalent - we checked the RMSE and it's not 0 at all. We did KLD on Q4_K_XL (the 155GB one other folks ship), and it's 96% top-1% agreement, so it's NOT lossless. llama.cpp does not have a native FP8 data-type, so BF16 is a must.

We could have named it MXFP4, but Q8_K_XL is the correct naming convention - we thought of BF16, but that's wrong. Also it's not all MXFP4, but MXFP4 + FP8

So when you see other people's quants at 155GB - this is not lossless and if they communicate that it's lossless - this is wrong. Converting FP8 to Q8_0 is not a lossless operation but lossy.

Full KLD table:

quant size GB mean KLD KLD 99% KLD 99.9% PPL top-1 %
UD-IQ1_S 82.5 0.66155 6.5134 11.3667 8.0932 72.30
UD-IQ1_M 86.9 0.58962 6.0813 10.4555 7.7384 74.08
UD-IQ2_XXS 90.9 0.48487 5.3472 9.7837 7.0909 76.60
UD-IQ2_M 90.9 0.48388 5.3461 9.7245 7.0894 76.56
UD-Q2_K_XL 96.8 0.40766 4.7839 9.0557 6.6782 78.57
UD-IQ3_XXS 104.2 0.30789 3.8147 7.7732 6.1972 81.93
UD-IQ3_S 116.1 0.26895 3.4845 7.3190 6.0266 83.06
UD-Q3_K_M 128.1 0.15734 2.1807 4.9571 5.7028 87.06
UD-Q3_K_XL 128.2 0.15751 2.1456 4.9708 5.7085 87.05
UD-IQ4_XS 136.7 0.11525 1.5929 3.9121 5.5522 88.79
UD-IQ4_NL 136.7 0.11525 1.5929 3.9121 5.5522 88.79
UD-Q4_K_XL 155.1 0.01324 0.1675 0.5512 5.3380 96.04
UD-Q8_K_XL 161.9 -0.00000 0.0000 0.0001 5.3322 100.00

Note UD-IQ3_XXS has been updated to reduce max KLD, and also some 3-bit quants are smaller in size now

Thanks for the clarification! In terms of activated parameters in Gbs what difference we are looking at between UD-Q4_K_XL and UD-Q8_K_XL? If it is +7Gb per token then probably I can live with 96% top1 agreement. It would be nice if you can state explicitly this parameter for all quantizations.

@perelmanych UD-Q4_K_XL and UD-Q8_K_XL are virtually equivalent, except UD-Q8_K_XL for attention, other non MoE layers are BF16, whilst UD-Q4_K_XL is in Q8_0

Sign up or log in to comment