Are other quants on the horizon?

#2
by InfernalDread - opened

Hello,

Thank you for being part of the few who are making quants of this model. Will you be creating smaller quants as well? Specifically within the range of 2-3 bit? The base model for this release is very resilient to quantization even down to the 2 bit range.

Thank you.

Yeah, admittedly I was only planning on uploading a Q4_K_M, Q6_K, and the BF16 GGUFs, under the guise that I'm sure other groups are doing more thorough releases. This was more intended as a quick compile since I know myself and others have been clamoring to at least try the 397B variant ha.

I also didn't use any imatrix calibration sets as I am new to this, so if I tried to pump out a lower quant it would for sure be a degradation in comparison to Bartowski's. While I could just use his imatrix dataset, that feels unfair to the amount of work he's put into that to get his own quants up and running.

All in all, I was treating the quant as a bit of a learning exercise for me with the added benefit of some real, usable files as the output. Actually makes me want to experiment with imatrix's some, see if I can't try to target generalized agentic use as a calibration set, and compare benchmark runs since the activations will be different. Might totally be a waste of time too lol, but that's a bit of the route I want to take.

Yeah, admittedly I was only planning on uploading a Q4_K_M, Q6_K, and the BF16 GGUFs, under the guise that I'm sure other groups are doing more thorough releases. This was more intended as a quick compile since I know myself and others have been clamoring to at least try the 397B variant ha.

I also didn't use any imatrix calibration sets as I am new to this, so if I tried to pump out a lower quant it would for sure be a degradation in comparison to Bartowski's. While I could just use his imatrix dataset, that feels unfair to the amount of work he's put into that to get his own quants up and running.

All in all, I was treating the quant as a bit of a learning exercise for me with the added benefit of some real, usable files as the output. Actually makes me want to experiment with imatrix's some, see if I can't try to target generalized agentic use as a calibration set, and compare benchmark runs since the activations will be different. Might totally be a waste of time too lol, but that's a bit of the route I want to take.

Not a problem! Thank you for still releasing the 2 quants anyways! I hope you have great success in the future!

Sign up or log in to comment