majentik commited on
Commit
14e1a9f
·
verified ·
1 Parent(s): 0f29d5d

docs: Tier 2 polish — variant matrix + quant trade-off

Browse files
Files changed (1) hide show
  1. README.md +42 -11
README.md CHANGED
@@ -2,21 +2,19 @@
2
  base_model: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
3
  library_name: mlx
4
  tags:
5
- - turboquant
6
- - kv-cache-quantization
7
- - nemotron
8
- - nvidia
9
- - mamba2
10
- - hybrid
11
- - quantized
12
- - mlx
13
- - 2bit
14
  license: other
15
  license_name: nvidia-open-model-license
16
  license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
17
  pipeline_tag: text-generation
18
- language:
19
- - en
20
  ---
21
 
22
  # Nemotron-3-Nano-4B - TurboQuant MLX 2-bit
@@ -95,3 +93,36 @@ This model requires approximately 1.2 GB of unified memory. Recommended hardware
95
  - [majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit) -- RotorQuant MLX 2-bit variant
96
  - [TurboQuant Paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
97
  - [MLX Framework](https://github.com/ml-explore/mlx)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  base_model: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
3
  library_name: mlx
4
  tags:
5
+ - turboquant
6
+ - kv-cache-quantization
7
+ - nemotron
8
+ - nvidia
9
+ - mamba2
10
+ - hybrid
11
+ - quantized
12
+ - mlx
13
+ - 2bit
14
  license: other
15
  license_name: nvidia-open-model-license
16
  license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
17
  pipeline_tag: text-generation
 
 
18
  ---
19
 
20
  # Nemotron-3-Nano-4B - TurboQuant MLX 2-bit
 
93
  - [majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit) -- RotorQuant MLX 2-bit variant
94
  - [TurboQuant Paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
95
  - [MLX Framework](https://github.com/ml-explore/mlx)
96
+
97
+ ## Quant trade-off (MLX lane)
98
+
99
+ | Bits | Approx size | Use case | Recommendation |
100
+ |---|---|---|---|
101
+ | **2-bit** | ~1.0 GB | Aggressive quantization | **Very low-RAM Macs** |
102
+ | 3-bit | ~1.4 GB | Lossy but small | Low-RAM Macs |
103
+ | 4-bit | ~1.7 GB | Balanced default | Recommended for most Macs |
104
+ | 5-bit | ~2.0 GB | Higher fidelity | Quality-sensitive |
105
+ | 6-bit | ~2.4 GB | Approaching FP16 quality | High-fidelity |
106
+ | 8-bit | ~3.0 GB | Near-lossless reference | Fidelity-critical work |
107
+
108
+ (Current variant — **2bit** — is bolded.)
109
+
110
+ ## Variants in this family
111
+
112
+ (Showing 13 sibling variants under `majentik/nemotron3-nano-4b-*`. The current variant — `TurboQuant-MLX-2bit` — is **bolded**.)
113
+
114
+ | Variant | Runtime | Approx size | Use case |
115
+ |---|---|---|---|
116
+ | [RotorQuant](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant) | runtime modifier | n/a | KV-cache root (weight-agnostic) |
117
+ | [RotorQuant-GGUF-IQ4_XS](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant-gguf-IQ4_XS) | llama.cpp | ~3.4 GB | Lossy 4-bit, low-RAM CPU/edge |
118
+ | [RotorQuant-GGUF-Q2_K](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant-gguf-Q2_K) | llama.cpp | ~2.4 GB | Lossy, low-RAM CPU/edge |
119
+ | [RotorQuant-GGUF-Q3_K_M](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant-gguf-Q3_K_M) | llama.cpp | ~3.1 GB | Smaller 3-bit, CPU-friendly |
120
+ | [RotorQuant-GGUF-Q4_K_M](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant-gguf-Q4_K_M) | llama.cpp | ~4.4 GB | Balanced default |
121
+ | [RotorQuant-MLX-2bit](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant-mlx-2bit) | mlx-lm | ~1.3 GB | Apple Silicon, smallest |
122
+ | [RotorQuant-MLX-4bit](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant-mlx-4bit) | mlx-lm | ~2.5 GB | Apple Silicon balanced |
123
+ | [RotorQuant-MLX-8bit](https://huggingface.co/majentik/nemotron3-nano-4b-rotorquant-mlx-8bit) | mlx-lm | ~4.7 GB | Apple Silicon reference |
124
+ | [TurboQuant](https://huggingface.co/majentik/nemotron3-nano-4b-turboquant) | runtime modifier | n/a | KV-cache root (weight-agnostic) |
125
+ | [TurboQuant-GGUF-Q4_K_M](https://huggingface.co/majentik/nemotron3-nano-4b-turboquant-gguf-Q4_K_M) | llama.cpp | ~4.4 GB | Balanced default |
126
+ | **TurboQuant-MLX-2bit** | mlx-lm | ~1.3 GB | Apple Silicon, smallest |
127
+ | [TurboQuant-MLX-4bit](https://huggingface.co/majentik/nemotron3-nano-4b-turboquant-mlx-4bit) | mlx-lm | ~2.5 GB | Apple Silicon balanced |
128
+ | [TurboQuant-MLX-8bit](https://huggingface.co/majentik/nemotron3-nano-4b-turboquant-mlx-8bit) | mlx-lm | ~4.7 GB | Apple Silicon reference |