--- language: - en - multilingual tags: - gemma - gemma-4 - qat - mxfp4 - blackwell - gguf - vision - multimodal license: apache-2.0 base_model: google/gemma-4-E4B-it-qat-q4_0-unquantized --- # Gemma 4 E4B Instruct - QAT + MXFP4 Hybrid GGUF **QAT-optimized weights preserved at Q4_0, overhead tensors quantized to MXFP4.** ## What Makes This Different This is a **hybrid quantization** of Google official QAT (Quantization-Aware Training) model. Instead of requantizing the Q4_0 weights (which breaks QAT benefits and vision quality), we: 1. **Kept all weight tensors at Q4_0** - attention, FFN, embeddings - exactly as Google trained them 2. **Quantized only the F32 norm/bias tensors to MXFP4** - these are the overhead tensors (layer norms, RMS norms, etc.) 3. **Used Google QAT mmproj** - the vision projector trained alongside the QAT model ### Why Standard MXFP4 from QAT Breaks Vision Google QAT model was specifically trained to be resilient to Q4_0 quantization patterns. The weight values learned during QAT compensate for Q4_0 rounding. When you requantize Q4_0 -> F32 -> MXFP4, a second round of quantization error is introduced that QAT training did **not** account for. Vision tokens flow through the same attention/FFN layers - precision loss disproportionately degrades vision. ### How the Hybrid Approach Works Using llama-quantize --tensor-type-file with --allow-requantize: ``` llama-quantize --allow-requantize --tensor-type-file keep_q4.txt input.gguf output.gguf MXFP4 ``` The tensor-type-file lists all Q4_0/Q4_K tensors to keep at their current type. When the quantizer sees cur_type == new_type, it copies the tensor data as-is - **zero precision loss**. Only the remaining F32 tensors are quantized to MXFP4. ## Usage ```bash # llama.cpp llama-server -m gemma-4-E4B-it-qat-mxfp4.gguf --mmproj mmproj-gemma-4-E4B-it-qat.gguf -ngl 99 ``` ## Source - **Base model**: [google/gemma-4-E4B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-E4B-it-qat-q4_0-unquantized) - **Quantized with**: llama.cpp build 537 (commit d2c6795) - **Chat template**: Native Gemma 4 (thinking enabled by default) - **Vision**: Full multimodal support via QAT mmproj ## Files | File | Description | |------|-------------| | gemma-4-E4B-it-qat-mxfp4.gguf | Q4_0 weights + MXFP4 norms | | mmproj-gemma-4-E4B-it-qat.gguf | QAT vision projector | ## License Apache 2.0 (same as base model)