Decoding speed ?
#1
by AliceThirty - opened
Looks interesting, I'd like to try if it has a decent speed. What's your hardware and your generation speed (token/s) with colibri ? I have a 9950x3D and a 192GB 5600MT/s for comparaison, and I can run GLM-5.2-UD-IQ1_M.gguf at 6 token/s with llama.cpp
It's slow enough to not be worth it for normal use tbh. I tried running Laguna, which is much smaller on some real agentic prompts and decode was a huge bottleneck. Consider all this experimental :)