WE need a Gemma 4 26B A4B version of this

#44
by CNWPlayer - opened

I can't run a dense model as fast as a MoE because my gpu is kinda poo, so I WANT A GEMMA 4 MoE VERSION OF THIS!

image

Waiting for 4 bit merge of branch to do this (tuning) on local hardware.
As of this writing sparse moes require training in 16 bit ; which with all the experts, context, and rank require roughly 70-80 GB of vram min.

oof, guessing you don't have that?

Sign up or log in to comment