File size: 9,191 Bytes
03ea283
1196d7d
 
 
03ea283
 
da268df
 
 
 
 
 
 
 
 
1196d7d
 
03ea283
 
b1b0de6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
03ea283
 
1196d7d
 
 
 
 
03ea283
da268df
 
 
 
 
 
 
 
03ea283
 
1196d7d
 
 
 
 
 
03ea283
 
 
b1b0de6
1196d7d
 
 
03ea283
1196d7d
 
 
 
 
 
 
 
 
 
 
 
 
 
03ea283
1196d7d
 
 
 
 
03ea283
 
1196d7d
 
 
 
 
03ea283
1196d7d
 
 
 
03ea283
 
1196d7d
 
 
 
 
 
 
 
 
 
 
 
 
03ea283
 
 
 
 
1196d7d
 
 
 
 
 
 
 
 
 
 
 
03ea283
2c54c38
03ea283
2c54c38
 
 
 
 
 
1196d7d
 
 
 
 
 
 
 
 
 
03ea283
1196d7d
03ea283
1196d7d
03ea283
1196d7d
 
 
 
 
03ea283
 
 
 
1196d7d
 
 
 
 
28f0927
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2c54c38
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
---
license: other
license_name: nvidia-open-model-license
license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
base_model: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
tags:
  - gguf
  - rotorquant
  - kv-cache-quantization
  - nemotron
  - nvidia
  - mamba2
  - hybrid
  - llama-cpp
  - quantized
library_name: gguf
pipeline_tag: text-generation
---

> [!TIP]
> **KV-cache quantization without any fork (recommended, 2026):** upstream
> llama.cpp/Ollama now cover this natively β€” use `-ctk q8_0 -ctv q8_0`
> (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
> `-ctk q4_0 -ctv q4_0` (~quarter memory, β‰ˆ7.6% perplexity increase). In
> Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
> K and V types symmetric to stay on the fast fused Flash-Attention path.
> Since April 2026, mainline llama.cpp also applies Hadamard rotation to
> KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
> which greatly improves low-bit KV quality (opt-out:
> `LLAMA_ATTN_ROT_DISABLE=1`).
>
> The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
> TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
> is unmaintained relative to mainline. It is NOT required to use this model.
<!-- kv-upstream-note -->

# Nemotron-3-Nano-4B-RotorQuant-GGUF-Q2_K

GGUF Q2_K weight-quantized variant of [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) optimised for use with **RotorQuant** KV cache compression via a dedicated llama.cpp fork.

> **Important:** RotorQuant KV cache types (`planar3`, `iso3`) are **not** available in upstream llama.cpp, standard Ollama, or LM Studio.
> They require a [specific llama.cpp fork](https://github.com/johndpope/llama-cpp-turboquant/tree/feature/planarquant-kv-cache).
> The GGUF file itself is a standard GGUF and works with any llama.cpp-compatible runtime using normal KV cache types (f16, q8_0, q4_0, etc.).

## Hardware compatibility

| Device | VRAM / RAM | Recommendation |
| --- | --- | --- |
| CPU host with β‰₯8 GB RAM | ~1.5 GB | works via llama.cpp; slower than GPU but no accelerator required |
| Apple Silicon (Metal) | ~1.7 GB | llama.cpp Metal backend; fast on M-series unified memory |
| NVIDIA GPU (partial offload) | split between GPU + RAM | offload as many layers as VRAM allows; rest on CPU |

## Overview

This model combines two independent compression techniques:

| Technique | What it does | Requirement |
|-----------|-------------|-------------|
| **GGUF Q2_K weight quantization** | Reduces model size from ~8 GB (BF16) to ~1.4 GB | Any llama.cpp-compatible runtime |
| **RotorQuant KV cache compression** β€” block-diagonal Clifford-algebra rotors for 3-bit KV cache (`--cache-type-k iso3 --cache-type-v iso3`) | Block-diagonal rotations / random rotation for compressed KV cache | [llama-cpp-turboquant fork](https://github.com/johndpope/llama-cpp-turboquant/tree/feature/planarquant-kv-cache) only |

## Quickstart

### Option A β€” RotorQuant KV cache (experimental fork β€” not required)

You must build from the RotorQuant-enabled llama.cpp fork:

```bash
# Clone and build the fork
git clone https://github.com/johndpope/llama-cpp-turboquant.git
cd llama-cpp-turboquant && git checkout feature/planarquant-kv-cache

# CUDA (Windows/Linux)
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build -j

# Metal (Apple Silicon)
cmake -B build -DGGML_METAL=ON -DGGML_METAL_EMBED_LIBRARY=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build -j

# Run with RotorQuant KV cache
./build/bin/llama-cli -m Nemotron-3-Nano-4B-RotorQuant-GGUF-Q2_K.gguf \
  --cache-type-k iso3 --cache-type-v iso3 \
  -ngl 99 -fa \
  -p "Explain quantum computing"

# Or run as a server
./build/bin/llama-server -m Nemotron-3-Nano-4B-RotorQuant-GGUF-Q2_K.gguf \
  --cache-type-k iso3 --cache-type-v iso3 \
  -ngl 99 -fa --jinja
```

### Option B β€” With standard llama.cpp / LM Studio / Ollama

The GGUF works as a normal quantised model. You won't get RotorQuant-specific KV cache benefits, but standard KV cache quantization (q8_0, q4_0) still reduces VRAM significantly.

**llama.cpp (upstream)**
```bash
llama-cli -m Nemotron-3-Nano-4B-RotorQuant-GGUF-Q2_K.gguf \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  -ngl 99 -fa \
  -p "Explain quantum computing"
```

**LM Studio**
1. Download the GGUF file and load in LM Studio.
2. Enable **Developer Mode** (Settings β†’ Developer).
3. In the model loader's advanced settings, set **Flash Attention** to ON.
4. Set **K Cache Quantization** and **V Cache Quantization** to `q8_0` (or `q4_0` for more aggressive VRAM savings).
5. Note: LM Studio does not currently support RotorQuant's `iso3` cache types. Track [this feature request](https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1719) for updates.

**Ollama**
```bash
# Standard Ollama does not support RotorQuant cache types.
# Use with default or q8_0 KV cache via OLLAMA_KV_CACHE_TYPE=q8_0
OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_FLASH_ATTENTION=1 ollama run majentik/Nemotron-3-Nano-4B-RotorQuant-GGUF-Q2_K
```

## Specifications

| Property | Value |
|----------|-------|
| Base Model | [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) |
| Architecture | Mamba-2 + Transformer hybrid (dense) |
| Parameters | 4B (dense hybrid) |
| Context Length | 262K |
| Weight Quantization | GGUF Q2_K (aggressive 2-bit, noticeable quality drop) |
| Original Size (BF16) | ~8 GB |
| Quantized File Size | ~1.4 GB |
| KV Cache (RotorQuant) | 3-bit via `--cache-type-k iso3 --cache-type-v iso3` (fork only) |
| KV Cache (standard) | q8_0, q4_0, f16, etc. (any llama.cpp runtime) |
| License | other |
| Modalities | Text only |
| Compatible Runtimes | llama.cpp, LM Studio, Ollama, koboldcpp |

## About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's **release labels**, not distinct
quantization algorithms β€” for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured. The KV-cache fork these
labels originally referred to is legacy; for KV-cache memory savings use the
upstream options described above (`-ctk/-ctv q8_0`, `OLLAMA_KV_CACHE_TYPE`).

## Current Status of RotorQuant in the Ecosystem

| Runtime | RotorQuant Support | Standard KV Quant |
|---------|---------------------|-------------------|
| llama.cpp (upstream) | ❌ Not merged | βœ… q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1 |
| llama-cpp-turboquant fork | βœ… planar3, iso3 | βœ… All standard types |
| LM Studio | ❌ [Requested](https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1719) | βœ… Via advanced settings |
| Ollama | ❌ Not supported | βœ… Via OLLAMA_KV_CACHE_TYPE |
| koboldcpp | ❌ Not supported | βœ… Standard types |

## Recommended Settings

For VRAM-constrained setups, standard q8_0 KV cache quantization already halves KV cache memory with negligible quality impact. Flash Attention should always be enabled β€” it is required for V cache quantization and improves memory efficiency regardless.

| VRAM | Suggested Configuration |
|------|------------------------|
| 24 GB (RTX 4090) | Q2_K + q8_0 KV cache + Flash Attention, 8K–16K context |
| 16 GB | Q2_K + q4_0 KV cache + Flash Attention, 4K–8K context |
| 48+ GB | Q2_K + f16 KV cache, full 32K+ context |

## See Also

- [RotorQuant GitHub](https://github.com/scrya-com/rotorquant)
- [llama-cpp-turboquant fork](https://github.com/johndpope/llama-cpp-turboquant/tree/feature/planarquant-kv-cache)
- [TurboQuant llama.cpp discussion](https://github.com/ggml-org/llama.cpp/discussions/20969)
- [TurboQuant paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
- [Base model: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16)
- [Nemotron-3-Nano-4B announcement](https://huggingface.co/blog/nvidia/nemotron-3-nano-4b)

## Quant trade-off (GGUF lane)

| Quant | Approx size | Use case | Recommendation |
|---|---|---|---|
| **Q2_K** | ~2.2 GB | Lossy, low-RAM CPU/edge | **Resource-constrained inference** |
| Q3_K_M | ~2.4 GB | Smaller-than-Q4, modest quality drop | Edge devices with ~16 GB RAM |
| IQ4_XS | ~2.1 GB | Importance-quant 4-bit, smaller than Q4_K_M | Best size/quality at 4-bit |
| Q4_K_M | ~3.0 GB | Balanced default | Recommended for most users |
| Q5_K_M | ~3.1 GB | Higher fidelity than Q4 | Quality-sensitive applications |
| Q6_K | ~3.6 GB | Approaching FP16 quality | High-fidelity CPU/edge |
| Q8_0 | ~4.1 GB | Near-lossless reference | Fidelity-critical work |
| MXFP4_MOE | ~2.2 GB | Microscaling FP4 (MoE-aware) | vLLM / transformers users |

(Current variant β€” **Q2_K** β€” is bolded.)

## Variants in this family

(Showing 13 sibling variants under `majentik/nemotron3-nano-4b-*`. The current variant β€” `RotorQuant-GGUF-Q2_K` β€” is **bolded**.)

| Variant | Runtime | Approx size | Use case |
|---|---|---|---|
| **RotorQuant-GGUF-Q2_K** | llama.cpp | ~2.4 GB | Lossy, low-RAM CPU/edge |