Hello! Did you try bigger size with mmap?

#2
by auf1r2 - opened

This model have ~50Gb of n-rgam weights, so potentilly we can use Q6 for Experts and stream n-grams from SSD theoretically without losing inference speed (at least not too much).

I can try it again I was getting oom on a few attempts I tried

Sign up or log in to comment