About the usage of this model for local companion (Flash Attention)

#4
by CARINB - opened

So i've been struggling to put this model to use, being unable to realibly fit it inside the 24gb of my 4090 with a decent context size, even by lending some vram(16gb) of my rtx 3090 my context stays around 30k (FP16, since gemma is said to loose quality fast, and i'm tired of gemma's 26b allucinations). Yet it seens i'm unable to apply FA (claude says it is unviable). So i'm asking community for help.

Sign up or log in to comment