majentik commited on
Commit
b1b0de6
·
verified ·
1 Parent(s): 28f0927

docs: upstream-first KV-cache guidance (q8_0/q4_0, mainline Hadamard rotation); fork demoted to experimental

Browse files
Files changed (1) hide show
  1. README.md +18 -1
README.md CHANGED
@@ -17,6 +17,23 @@ library_name: gguf
17
  pipeline_tag: text-generation
18
  ---
19
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
  # Nemotron-3-Nano-4B-RotorQuant-GGUF-Q2_K
21
 
22
  GGUF Q2_K weight-quantized variant of [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) optimised for use with **RotorQuant** KV cache compression via a dedicated llama.cpp fork.
@@ -44,7 +61,7 @@ This model combines two independent compression techniques:
44
 
45
  ## Quickstart
46
 
47
- ### Option A — With RotorQuant KV cache (fork required)
48
 
49
  You must build from the RotorQuant-enabled llama.cpp fork:
50
 
 
17
  pipeline_tag: text-generation
18
  ---
19
 
20
+ > [!TIP]
21
+ > **KV-cache quantization without any fork (recommended, 2026):** upstream
22
+ > llama.cpp/Ollama now cover this natively — use `-ctk q8_0 -ctv q8_0`
23
+ > (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
24
+ > `-ctk q4_0 -ctv q4_0` (~quarter memory, ≈7.6% perplexity increase). In
25
+ > Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
26
+ > K and V types symmetric to stay on the fast fused Flash-Attention path.
27
+ > Since April 2026, mainline llama.cpp also applies Hadamard rotation to
28
+ > KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
29
+ > which greatly improves low-bit KV quality (opt-out:
30
+ > `LLAMA_ATTN_ROT_DISABLE=1`).
31
+ >
32
+ > The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
33
+ > TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
34
+ > is unmaintained relative to mainline. It is NOT required to use this model.
35
+ <!-- kv-upstream-note -->
36
+
37
  # Nemotron-3-Nano-4B-RotorQuant-GGUF-Q2_K
38
 
39
  GGUF Q2_K weight-quantized variant of [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) optimised for use with **RotorQuant** KV cache compression via a dedicated llama.cpp fork.
 
61
 
62
  ## Quickstart
63
 
64
+ ### Option A — RotorQuant KV cache (experimental fork — not required)
65
 
66
  You must build from the RotorQuant-enabled llama.cpp fork:
67