gemma-4-12B-it — TurboQuant MLX NVFP4

google/gemma-4-12B-it @ 12ace6d648d72bd41519e140f1185f34d38c7e3d quantized pack, published as majentik/gemma-4-12B-it-TurboQuant-MLX-NVFP4.

Method

MLX quantization via mlx-vlm 0.6.3 (NVFP4, group_size 16); vision + audio towers retained in BF16 (not quantized).

Release line

Released under the TurboQuant line. RotorQuant and TurboQuant are this project's release labels for this pack, not distinct quantization algorithms — both brand repos for a given tier carry byte-identical weights, produced once and published under two names. No brand-specific speedup is claimed or measured for either label.

Modality

This pack is image-text-to-text capable: the vision and audio towers ship in BF16 alongside the quantized text tower, so image (and audio) inputs are supported end to end via mlx-vlm.

Vision fidelity note: for this instruct variant, vision has only been validated at 6-bit and 8-bit. Vision fidelity degrades below 6-bit; use 6-bit or 8-bit for image tasks.

License

Governed by the Gemma Terms of Use. See the upstream repo for the full license text.

Downloads last month
82
Safetensors
Model size
3B params
Tensor type
U8
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for majentik/gemma-4-12B-it-TurboQuant-MLX-NVFP4

Quantized
(288)
this model