majentik commited on
Commit
daf21e2
·
verified ·
1 Parent(s): 14e1a9f

docs: upstream-first KV-cache guidance (q8_0/q4_0, mainline Hadamard rotation); fork demoted to experimental

Browse files
Files changed (1) hide show
  1. README.md +17 -0
README.md CHANGED
@@ -17,6 +17,23 @@ license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-l
17
  pipeline_tag: text-generation
18
  ---
19
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
  # Nemotron-3-Nano-4B - TurboQuant MLX 2-bit
21
 
22
  **2-bit weight-quantized MLX version** of [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) with TurboQuant KV-cache quantization. Optimized for Apple Silicon inference via the [MLX](https://github.com/ml-explore/mlx) framework. Maximum compression for running on memory-constrained devices. The dense hybrid Mamba-2 + Attention architecture supports up to 262K context length.
 
17
  pipeline_tag: text-generation
18
  ---
19
 
20
+ > [!TIP]
21
+ > **KV-cache quantization without any fork (recommended, 2026):** upstream
22
+ > llama.cpp/Ollama now cover this natively — use `-ctk q8_0 -ctv q8_0`
23
+ > (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
24
+ > `-ctk q4_0 -ctv q4_0` (~quarter memory, ≈7.6% perplexity increase). In
25
+ > Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
26
+ > K and V types symmetric to stay on the fast fused Flash-Attention path.
27
+ > Since April 2026, mainline llama.cpp also applies Hadamard rotation to
28
+ > KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
29
+ > which greatly improves low-bit KV quality (opt-out:
30
+ > `LLAMA_ATTN_ROT_DISABLE=1`).
31
+ >
32
+ > The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
33
+ > TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
34
+ > is unmaintained relative to mainline. It is NOT required to use this model.
35
+ <!-- kv-upstream-note -->
36
+
37
  # Nemotron-3-Nano-4B - TurboQuant MLX 2-bit
38
 
39
  **2-bit weight-quantized MLX version** of [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) with TurboQuant KV-cache quantization. Optimized for Apple Silicon inference via the [MLX](https://github.com/ml-explore/mlx) framework. Maximum compression for running on memory-constrained devices. The dense hybrid Mamba-2 + Attention architecture supports up to 262K context length.