INT8 W8A8 Quant for vLLM?

#51
by bSun0000 - opened

Does anyone have plans to make it? INT8 weights + INT8 activation.
This should be the best (performance-wise) quant to run on Ampere cards, utilizing tensor cores, isnt?
For some reason, W4A16 & W8A16 quants seems to be more popular, but why?

Sign up or log in to comment