How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "majentik/gemma-4-12B-it-TurboQuant-MLX-NVFP4"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default majentik/gemma-4-12B-it-TurboQuant-MLX-NVFP4
Run Hermes
hermes
Quick Links

gemma-4-12B-it — TurboQuant MLX NVFP4

google/gemma-4-12B-it @ 12ace6d648d72bd41519e140f1185f34d38c7e3d quantized pack, published as majentik/gemma-4-12B-it-TurboQuant-MLX-NVFP4.

Method

MLX quantization via mlx-vlm 0.6.3 (NVFP4, group_size 16); vision + audio towers retained in BF16 (not quantized).

Release line

Released under the TurboQuant line. RotorQuant and TurboQuant are this project's release labels for this pack, not distinct quantization algorithms — both brand repos for a given tier carry byte-identical weights, produced once and published under two names. No brand-specific speedup is claimed or measured for either label.

Modality

This pack is image-text-to-text capable: the vision and audio towers ship in BF16 alongside the quantized text tower, so image (and audio) inputs are supported end to end via mlx-vlm.

Vision fidelity note: for this instruct variant, vision has only been validated at 6-bit and 8-bit. Vision fidelity degrades below 6-bit; use 6-bit or 8-bit for image tasks.

License

Governed by the Gemma Terms of Use. See the upstream repo for the full license text.

Downloads last month
82
Safetensors
Model size
3B params
Tensor type
U8
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for majentik/gemma-4-12B-it-TurboQuant-MLX-NVFP4

Quantized
(288)
this model