Qwen3.6 27B family evaluation

#3
by tcclaviger - opened

Standby, A large family wide evaluation landing soon...

I have hit a ... strange quirk, assessing.. The BF16 model is showing regression from the GGUF Q8_0 model that I unpacked into BF16 and tested, was this QAT/QLORA merged, might explain part of it?

Both are showing an anomalous behavior in a NIAH test I've not observed in 27B variants before.

Implementing a feature in vllm that will help me dissect it, overall there IS an improvement that is incontrovertible vs the base model, but I'm trying to understand what's going on with the anomaly before I drop results.

I am running overnight a test for the BF16 made from source to establish a baseline. This will take about 10-12 hours.

Where it stands now can be viewed here, I am working to resolve the mentioned anomalies above, when done I will rerun the whole family to give each an honest shake down. In short, I'm implementing DRY in vllm to clamp thinking repetition across the whole family so they get a fair capability shakedown and not a "does this model loop" shakedown, because full transparency...every single 27B will thinking loop if stressed hard enough.

https://blog.robai.net/27bevals/

I looked at the metrics, they make complete sense.

I see higher numbers with qx86-hi and mxfp8 that are consistently higher than BF16 in previous models.

This is normal, and I rarely see a BF16 outperform a tuned quant. The GGUF XL mixed quants and David's NEO raise those metrics even further.

On MLX I see the same effect on qx86-hi and qx64-hi, here is for example the Qwen3.6-27B-Architect-DS9-Polaris2

         arc   arc/e boolq hswag obkqa piqa  wino
bf16     0.692,0.863,0.911
mxfp8    0.699,0.871,0.910
q8-hi    0.694,0.865,0.910
qx86-hi  0.688,0.862,0.910
qx64-hi  0.700,0.862,0.907
mxfp4    0.694,0.872,0.909

Quant    Perplexity      Peak Memory   Tokens/sec
bf16     3.898 ± 0.025   60.75 GB      226
q8-hi    3.895 ± 0.025   37.26 GB      215
mxfp8    3.921 ± 0.025   34.74 GB      218
qx86-hi  3.898 ± 0.025   32.36 GB      218
qx64-hi  3.918 ± 0.025   25.64 GB      217
mxfp4    3.999 ± 0.025   21.30 GB      225

If you compare now how the BF16 and F32 is quanting down, mxfp4 is the only one that did not change at all.

From F32, the lower quants are much better, while mxfp8 suffers, as expected. I did not benchmark the q8-hi from F32 to compare it to BF16, that would be an interesting test.

The F32 source is available.

Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
mxfp8     0.711,0.879,0.910,0.790,0.514,0.823,0.763
qx86-hi   0.696,0.876,0.912,0.791,0.518,0.824,0.760
qx64-hi   0.702,0.873,0.909,0.794,0.514,0.822,0.750
mxfp4     0.701,0.873,0.909,0.786,0.488,0.813,0.759
nvfp4     0.701,0.866,0.907
Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     3.783 ± 0.023   34.74 GB      203
qx86-hi   3.735 ± 0.023   33.25 GB      183
qx64-hi   3.747 ± 0.023   27.03 GB      194
mxfp4     3.854 ± 0.024   21.30 GB      197


Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-F32
mxfp8     0.709,0.880,0.909
qx64-hi   0.706,0.873,0.908
nvfp4     0.704,0.868,0.908
mxfp4     0.701,0.873,0.909

Also, while you are at it, there is another one

https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Tess

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.712,0.879,0.911,0.792,0.508,0.823,0.764
qx86-hi   0.701,0.877,0.911,0.794,0.518,0.823,0.758
qx64-hi   0.706,0.873,0.909,0.795,0.512,0.823,0.752
mxfp4     0.706,0.873,0.910,0.790,0.496,0.817,0.761
1M
mxfp8     0.703,0.877,0.909
qx86-hi   0.702,0.873,0.911,0.793,0.508,0.824,0.765
qx64-hi   0.706,0.873,0.909,0.795,0.512,0.823,0.752
mxfp4     0.701,0.874,0.912,0.789,0.500,0.817,0.759

Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     3.797 ± 0.024   34.74 GB      183
qx86-hi   3.756 ± 0.023   33.25 GB      170
qx64-hi   3.765 ± 0.023   27.03 GB      181
mxfp4     3.876 ± 0.024   21.26 GB      176
1M
mxfp8     3.803 ± 0.024   34.70 GB      169
qx86-hi   3.758 ± 0.023   33.21 GB      175
qx64-hi   3.769 ± 0.023   26.99 GB      172
mxfp4     3.876 ± 0.024   21.26 GB      176

That validates some other model and quant testing I've done, it seems the distribution snapping impact of quantization seems to make certain behaviors have a higher probability in a way that is represented as better performance in testing.

I've been trying to figure out why my custom 4 bit always leads to better performance than 16 bit in certain tests so, while I might not fully understand the exact mechanism, its good to know others are also observing the effect. Thanks for the responses!

Sign up or log in to comment