File size: 5,706 Bytes
fae218e
 
 
 
14e1a9f
 
 
 
 
 
 
 
 
fae218e
 
 
 
 
 
daf21e2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fae218e
 
e3e8b9e
fae218e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e3e8b9e
fae218e
e3e8b9e
 
 
 
 
 
fae218e
 
 
 
 
 
e3e8b9e
fae218e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14e1a9f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e3e8b9e
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
---
base_model: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
library_name: mlx
tags:
  - turboquant
  - kv-cache-quantization
  - nemotron
  - nvidia
  - mamba2
  - hybrid
  - quantized
  - mlx
  - 2bit
license: other
license_name: nvidia-open-model-license
license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
pipeline_tag: text-generation
---

> [!TIP]
> **KV-cache quantization without any fork (recommended, 2026):** upstream
> llama.cpp/Ollama now cover this natively — use `-ctk q8_0 -ctv q8_0`
> (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
> `-ctk q4_0 -ctv q4_0` (~quarter memory, ≈7.6% perplexity increase). In
> Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
> K and V types symmetric to stay on the fast fused Flash-Attention path.
> Since April 2026, mainline llama.cpp also applies Hadamard rotation to
> KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
> which greatly improves low-bit KV quality (opt-out:
> `LLAMA_ATTN_ROT_DISABLE=1`).
>
> The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
> TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
> is unmaintained relative to mainline. It is NOT required to use this model.
<!-- kv-upstream-note -->

# Nemotron-3-Nano-4B - TurboQuant MLX 2-bit

**2-bit weight-quantized MLX version** of [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) with the legacy TurboQuant KV-cache fork (superseded by upstream llama.cpp KV options). Optimized for Apple Silicon inference via the [MLX](https://github.com/ml-explore/mlx) framework. Maximum compression for running on memory-constrained devices. The dense hybrid Mamba-2 + Attention architecture supports up to 262K context length.

Approximate model size: **~1.2 GB**

## Model Specifications

| Property | Value |
|---|---|
| **Base Model** | [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) |
| **Parameters** | 4 billion (dense) |
| **Architecture** | Hybrid Mamba-2 + Attention (dense) |
| **Context Length** | 262,144 tokens (262K) |
| **License** | NVIDIA Open Model License (commercial use OK) |
| **Weight Quantization** | 2-bit (~1.2 GB) |
| **KV-Cache Quantization** | TurboQuant |
| **Framework** | MLX (Apple Silicon) |

## Quickstart

```python
from mlx_lm import load, generate
from turboquant import TurboQuantCache

model, tokenizer = load("majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-2bit")

prompt = "Explain the theory of relativity."
response = generate(model, tokenizer, prompt=prompt, max_tokens=512)
print(response)
```

## About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's **release labels**, not distinct
quantization algorithms — for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured. The KV-cache fork these
labels originally referred to is legacy; for KV-cache memory savings use the
upstream options described above (`-ctk/-ctv q8_0`, `OLLAMA_KV_CACHE_TYPE`).

## KV-Cache Quantization Comparison

| Method | Prefill Speed | Decode Speed | Memory Savings | Reference |
|---|---|---|---|---|
| **TurboQuant** | 1x (baseline) | 1x (baseline) | High | [arXiv: 2504.19874](https://arxiv.org/abs/2504.19874) |


## Memory Estimates (Nemotron-3-Nano-4B)

| Precision | Approximate Size | MLX Variant |
|---|---|---|
| BF16 (original) | ~8 GB | -- |
| 8-bit quantized | ~4 GB | [TurboQuant-MLX-8bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-8bit) |
| 4-bit quantized | ~2.3 GB | [TurboQuant-MLX-4bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-4bit) |
| **2-bit quantized** | **~1.2 GB** | **This model** |

## Hardware Requirements

This model requires approximately 1.2 GB of unified memory. Recommended hardware:
- Apple M1 (8 GB+)
- Apple M2 (8 GB+)
- Apple M3 (8 GB+)
- Apple M4 (8 GB+)
- Any Apple Silicon Mac with 8 GB+ unified memory

## See Also

- [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) -- Base model
- [majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-8bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-8bit) -- MLX 8-bit variant
- [majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-4bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-4bit) -- MLX 4-bit variant
- [majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit) -- RotorQuant MLX 2-bit variant
- [TurboQuant Paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
- [MLX Framework](https://github.com/ml-explore/mlx)

## Quant trade-off (MLX lane)

| Bits | Approx size | Use case | Recommendation |
|---|---|---|---|
| **2-bit** | ~1.0 GB | Aggressive quantization | **Very low-RAM Macs** |
| 3-bit | ~1.4 GB | Lossy but small | Low-RAM Macs |
| 4-bit | ~1.7 GB | Balanced default | Recommended for most Macs |
| 5-bit | ~2.0 GB | Higher fidelity | Quality-sensitive |
| 6-bit | ~2.4 GB | Approaching FP16 quality | High-fidelity |
| 8-bit | ~3.0 GB | Near-lossless reference | Fidelity-critical work |

(Current variant — **2bit** — is bolded.)

## Variants in this family

(Showing 13 sibling variants under `majentik/nemotron3-nano-4b-*`. The current variant — `TurboQuant-MLX-2bit` — is **bolded**.)

| Variant | Runtime | Approx size | Use case |
|---|---|---|---|
| **TurboQuant-MLX-2bit** | mlx-lm | ~1.3 GB | Apple Silicon, smallest |