Qwen3.5-9B-GGUF-MoQ / README.md
w-ahmad's picture
Update README.md
93149d6 verified
|
Raw
History Blame
7.11 kB
---
language:
- en
library_name: gguf
tags:
- MoQ
- mixture-of-quants
- GGUF
- QWEN
- quantization
base_model:
- Qwen/Qwen3.5-9B
license: mit
pipeline_tag: text-generation
---
# πŸš€ MoQ: Mixture of Quants
>MoQ (Mixture of Quants) is a smart way to shrink AI models without losing their "brainpower." Unlike old methods that treat every part of the model the same, MoQ identifies the most important parts and keeps them high-quality, while heavily compressing the rest to save space.****Stop settling for uniform bitrates.** Standard quantization is a relic of the past, treating vital cognitive weights the same as redundant noise. **MoQ (Mixture of Quants)** is a surgical evolution in model compression. By deploying an **Empirical Per-Tensor Analysis**, MoQ identifies the "High-Intelligence" tensors that drive reasoning and shields them with high-bit precision, while crushing redundant weights into extreme efficiency.
---
The result? A model that punches significantly above its weight class.
---
## πŸ“₯ Available Quants
| Folder Link | BPW | Total Size | Description |
| :--- | :---: | :---: | :--- |
[πŸ“‚ **BF-16**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/BF-16/Qwen3.5-9B-MoQ-16.gguf) | **16** | **~16.69 GB**
| [πŸ“‚ **FP-16**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/FP-16/Qwen3.5-9B-MoQ-16.gguf) | **16** | **~16.69 GB**
| [πŸ“‚ **MoQ-2.10**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-2.10/Qwen3.5-9B-MoQ-2.10.gguf) | **2.10** | **~2.23 GB**
| [πŸ“‚ **MoQ-2.55**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-2.55/Qwen3.5-9B-MoQ-2.55.gguf) | **2.55** | **~2.68 GB**
| [πŸ“‚ **MoQ-2.80**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-2.80/Qwen3.5-9B-MoQ-2.80.gguf) | **2.80** | **~2.90 GB**
| [πŸ“‚ **MoQ-2.85**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-2.85/Qwen3.5-9B-MoQ-2.85.gguf) | **2.85** | **~3.00 GB**
| [πŸ“‚ **MoQ-2.90**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-2.90/Qwen3.5-9B-MoQ-2.90.gguf) | **2.90** | **~3.03 GB**
| [πŸ“‚ **MoQ-3.10**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-3.10/Qwen3.5-9B-MoQ-3.10.gguf) | **3.10** | **~3.24 GB**
| [πŸ“‚ **MoQ-3.25**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-3.25/Qwen3.5-9B-MoQ-3.25.gguf) | **3.25** | **~3.41 GB**
| [πŸ“‚ **MoQ-3.30**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-3.30/Qwen3.5-9B-MoQ-3.30.gguf) | **3.30** | **~3.44 GB**
| [πŸ“‚ **MoQ-3.65**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-3.65/Qwen3.5-9B-MoQ-3.65.gguf) | **3.65** | **~3.79 GB**
| [πŸ“‚ **MoQ-3.75**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-3.75/Qwen3.5-9B-MoQ-3.75.gguf) | **3.75** | **~3.92 GB**
| [πŸ“‚ **MoQ-4.00**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-4.00/Qwen3.5-9B-MoQ-4.00.gguf) | **4.00** | **~4.20 GB**
| [πŸ“‚ **MoQ-4.25**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-4.25/Qwen3.5-9B-MoQ-4.25.gguf) | **4.25** | **~4.43 GB**
| [πŸ“‚ **MoQ-4.65**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-4.65/Qwen3.5-9B-MoQ-4.65.gguf) | **4.65** | **~4.87 GB**
| [πŸ“‚ **MoQ-4.85**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-4.85/Qwen3.5-9B-MoQ-4.85.gguf) | **4.85** | **~5.05 GB**
| [πŸ“‚ **MoQ-5.00**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-5.00/Qwen3.5-9B-MoQ-5.00.gguf) | **5.00** | **~5.24 GB**
| [πŸ“‚ **MoQ-5.65**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-5.65/Qwen3.5-9B-MoQ-5.65.gguf) | **5.65** | **~5.91 GB**
| [πŸ“‚ **MoQ-6.35**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-6.35/Qwen3.5-9B-MoQ-6.35.gguf) | **6.35** | **~6.61 GB**
| [πŸ“‚ **MoQ-7.80**](https://huggingface.co/w-ahmad/Qwen3.5-9B-GGUF-MoQ/blob/main/MoQ-7.80/Qwen3.5-9B-MoQ-7.80.gguf) | **7.80** | **~8.17 GB**
---
## πŸ“Š MoQ vs. Unsloth: Performance Comparison
The following benchmarks compare **MoQ 4.85** against **Unsloth Dynamic Quants**.
*Note: Lower KLD (Kullback–Leibler Divergence) indicates higher fidelity to the original model.*
### πŸ“‰ Key Divergence Metrics (Lower is Better)
#### **Mean KLD**
Average divergence across all layers. MoQ maintains a significantly lower average error profile.
![Mean_KLD](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/vongwACIdL5vim_VLwcIn.png)
#### **Maximum KLD**
The "worst-case" divergence point. MoQ 4.84 effectively eliminates the extreme divergence spikes seen in standard dynamic quants.
![Maximum_KLD](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/ThF0qh5xS45y1iwVeWuH7.png)
#### **RMS Ξ”p**
Root Mean Square change in probabilities. This measures the stability of the model's confidence.
![RMS_Ξ”p](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/0pglHo80A6IjJRjPM0Zjs.png)
---
### πŸ“ˆ Precision Percentiles
These graphs demonstrate MoQ's ability to maintain stability even within the most sensitive portions of the architecture.
#### **95.0% Percentile KLD**
![95.0pct_KLD](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/G9jJMxUgBUR0unkU3fX6q.png)
#### **99.0% Percentile KLD**
![99.0pct_KLD](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/HhgKNfcXXyMio5z84vIGI.png)
#### **99.9% Percentile KLD**
![99.9pct_KLD](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/ZqlPEtkIlCRH_0HL3-nlm.png)
---
### 🎯 Token Match (Same Top-P)
This metric tracks how often the quantized model chooses the **exact same top token** as the original high-precision model.
![Same_top_p](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/uSciX-oKfWFTwkPxOvvPO.png)
---
## 🧠 The MoQ Edge
MoQ optimizes the architecture for the **Pareto frontier** of memory and performance.
* **Dynamic Bitrate Allocation:** No more "one-size-fits-all." MoQ assigns precision where it actually matters.
* **Cognitive Preservation:** Massive VRAM savings with near-zero degradation in logic and coherence.
* **Next-Gen Efficiency:** Fits "Large" model intelligence into "Small" model hardware.
### πŸ“ Weight Sensitivity Heatmap
Lighter regions represent mission-critical tensors preserved at higher precision.
![importance_heatmap](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/xrFfZoqhPquwX9EUl3D5O.png)
### πŸ“ˆ Importance Distribution
The histogram shows the importance scores used to mathematically determine the optimal quant for each tensor.
![importance_histogram](https://cdn-uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/FgJJohu80wXmZqIH4ifwB.png)
##
Follow me on Linkedin
linkedin.com/in/waleed-ahmad-8a3166403
If MoQ does not perform well, email me :
waleedahmad.1a10@gmail.com
## πŸ›  Usage & Deployment
```bash
./llama-cli -m Qwen3.5-9B-MoQ-4.85.gguf -p "The future of efficient AI is..."