--- license: apache-2.0 license_link: https://ai.google.dev/gemma/docs/gemma_4_license library_name: transformers pipeline_tag: text-generation base_model: google/gemma-4-12B-it tags: - gemma4 - gemma4-unified - fp8 - modelopt - quantized - vllm --- # Gemma 4 12B-it — Text FP8 (ModelOpt) FP8-quantized text tower of [`google/gemma-4-12B-it`](https://huggingface.co/google/gemma-4-12B-it), the unified (encoder-free multimodal) Gemma 4 12B model. The linear layers of the language model are quantized to FP8 (E4M3) with per-tensor static scales calibrated offline via NVIDIA [ModelOpt](https://github.com/NVIDIA/TensorRT-Model-Optimizer); `lm_head` and the (tied) embeddings stay in BF16. This checkpoint is **text-only**: the vision/audio encoder weights are not included. The config still advertises `Gemma4UnifiedForConditionalGeneration` and the weights live under `model.language_model.*`, which keeps the layer namespace compatible with the paired MTP drafter for speculative decoding (see below). Serve it with `--limit-mm-per-prompt '{"image": 0, "audio": 0}'` so the absent multimodal encoder is never invoked. Produced by the same pipeline as [`bahadirakdemir/gemma-4-31B-it-text-fp8`](https://huggingface.co/bahadirakdemir/gemma-4-31B-it-text-fp8). ## Requirements This is the **unified** Gemma 4 architecture (`model_type: gemma4_unified`), which is newer than the classic `gemma4` (e.g. 31B). You need: - **transformers ≥ 5.10.0** (when `gemma4_unified` was added) - **vLLM with `gemma4_unified` support** — at the time of writing this is on the `main` branch / nightly (`uv pip install -U vllm --pre`), not yet in a tagged stable release (≤ 0.22.0). It will be in the next stable release. ## Usage with vLLM ```bash vllm serve bahadirakdemir/gemma-4-12B-it-text-fp8 \ --quantization modelopt \ --max-model-len 8192 \ --max-num-batched-tokens 8192 \ --gpu-memory-utilization 0.5 \ --limit-mm-per-prompt '{"image": 0, "audio": 0}' ``` For speculative decoding, pair it with the matching FP8 MTP drafter [`bahadirakdemir/gemma-4-12B-it-assistant-fp8`](https://huggingface.co/bahadirakdemir/gemma-4-12B-it-assistant-fp8): ```bash vllm serve bahadirakdemir/gemma-4-12B-it-text-fp8 \ --quantization modelopt \ --max-model-len 8192 \ --max-num-batched-tokens 8192 \ --gpu-memory-utilization 0.5 \ --limit-mm-per-prompt '{"image": 0, "audio": 0}' \ --speculative-config '{"model": "bahadirakdemir/gemma-4-12B-it-assistant-fp8", "num_speculative_tokens": 4}' ``` Tested with `vllm/vllm-openai:gemma4-0505-arm64-cu130` on NVIDIA GB10. ## Quantization details | | | |---|---| | Method | ModelOpt FP8 PTQ (E4M3, per-tensor static scales) | | Quantized | language-model linears (attention + MLP projections) | | Kept in BF16 | `lm_head`, tied embeddings, all norms | | Calibration | 32 instruct-style prompts, max length 1024 | License: Apache 2.0, inherited from upstream Gemma 4 — see [the Gemma 4 license](https://ai.google.dev/gemma/docs/gemma_4_license).