majentik commited on
Commit
fae218e
·
verified ·
1 Parent(s): de51832

Add model card

Browse files
Files changed (1) hide show
  1. README.md +95 -0
README.md ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
3
+ library_name: mlx
4
+ tags:
5
+ - turboquant
6
+ - kv-cache-quantization
7
+ - nemotron
8
+ - nvidia
9
+ - mamba2
10
+ - hybrid
11
+ - quantized
12
+ - mlx
13
+ - 2bit
14
+ license: other
15
+ license_name: nvidia-open-model-license
16
+ license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
17
+ pipeline_tag: text-generation
18
+ ---
19
+
20
+ # Nemotron-3-Nano-4B - TurboQuant MLX 2-bit
21
+
22
+ **2-bit weight-quantized MLX version** of [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) with TurboQuant KV-cache quantization. Optimized for Apple Silicon inference via the [MLX](https://github.com/ml-explore/mlx) framework. Maximum compression for running on memory-constrained devices. The dense hybrid Mamba-2 + Attention architecture supports up to 262K context length.
23
+
24
+ Approximate model size: **~1.2 GB**
25
+
26
+ ## Model Specifications
27
+
28
+ | Property | Value |
29
+ |---|---|
30
+ | **Base Model** | [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) |
31
+ | **Parameters** | 4 billion (dense) |
32
+ | **Architecture** | Hybrid Mamba-2 + Attention (dense) |
33
+ | **Context Length** | 262,144 tokens (262K) |
34
+ | **License** | NVIDIA Open Model License (commercial use OK) |
35
+ | **Weight Quantization** | 2-bit (~1.2 GB) |
36
+ | **KV-Cache Quantization** | TurboQuant |
37
+ | **Framework** | MLX (Apple Silicon) |
38
+
39
+ ## Quickstart
40
+
41
+ ```python
42
+ from mlx_lm import load, generate
43
+ from turboquant import TurboQuantCache
44
+
45
+ model, tokenizer = load("majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-2bit")
46
+
47
+ prompt = "Explain the theory of relativity."
48
+ response = generate(model, tokenizer, prompt=prompt, max_tokens=512)
49
+ print(response)
50
+ ```
51
+
52
+ ## What is TurboQuant?
53
+
54
+ TurboQuant ([arXiv: 2504.19874](https://arxiv.org/abs/2504.19874)) is a KV-cache quantization technique that compresses the key-value cache used during autoregressive generation. Combined with 2-bit weight quantization in MLX, this provides a dual compression strategy: smaller model weights plus compressed KV cache for efficient long-context generation.
55
+
56
+ Key benefits:
57
+ - **No weight modification** -- model weights stay at original precision
58
+ - **Reduced inference memory** -- KV cache is compressed significantly
59
+ - **Longer context windows** -- fit more tokens in the same GPU memory
60
+ - **Minimal quality loss** -- carefully designed quantization preserves generation quality
61
+
62
+ ## KV-Cache Quantization Comparison
63
+
64
+ | Method | Prefill Speed | Decode Speed | Memory Savings | Reference |
65
+ |---|---|---|---|---|
66
+ | **TurboQuant** | 1x (baseline) | 1x (baseline) | High | [arXiv: 2504.19874](https://arxiv.org/abs/2504.19874) |
67
+ | **RotorQuant** | **5.3x faster** | **28% faster** | High | [GitHub](https://github.com/scrya-com/rotorquant) |
68
+
69
+ ## Memory Estimates (Nemotron-3-Nano-4B)
70
+
71
+ | Precision | Approximate Size | MLX Variant |
72
+ |---|---|---|
73
+ | BF16 (original) | ~8 GB | -- |
74
+ | 8-bit quantized | ~4 GB | [TurboQuant-MLX-8bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-8bit) |
75
+ | 4-bit quantized | ~2.3 GB | [TurboQuant-MLX-4bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-4bit) |
76
+ | **2-bit quantized** | **~1.2 GB** | **This model** |
77
+
78
+ ## Hardware Requirements
79
+
80
+ This model requires approximately 1.2 GB of unified memory. Recommended hardware:
81
+ - Apple M1 (8 GB+)
82
+ - Apple M2 (8 GB+)
83
+ - Apple M3 (8 GB+)
84
+ - Apple M4 (8 GB+)
85
+ - Any Apple Silicon Mac with 8 GB+ unified memory
86
+
87
+ ## See Also
88
+
89
+ - [nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16) -- Base model
90
+ - [majentik/Nemotron-3-Nano-4B-TurboQuant](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant) -- TurboQuant KV-cache only (transformers)
91
+ - [majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-8bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-8bit) -- MLX 8-bit variant
92
+ - [majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-4bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-TurboQuant-MLX-4bit) -- MLX 4-bit variant
93
+ - [majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit](https://huggingface.co/majentik/Nemotron-3-Nano-4B-RotorQuant-MLX-2bit) -- RotorQuant MLX 2-bit variant
94
+ - [TurboQuant Paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
95
+ - [MLX Framework](https://github.com/ml-explore/mlx)