michaelw9999 commited on
Commit
287d5c2
·
verified ·
1 Parent(s): 1970b79

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +5 -0
README.md CHANGED
@@ -22,6 +22,11 @@ Recently re-quantized with the latest version of <A HREF="https://github.com/mic
22
  It applies my RSF scale fitting technique to the Q_K quants and uses a different tensor mix layout compared to previously.
23
  6-June-2026: changed the MTP tensors to be NVFP4, ~130tk/s tg is now possible in some configurations on 5090.
24
  </p>
 
 
 
 
 
25
  <p>Feedback to improve this is appreciated.</p>
26
  Results below on RTX 5090 with Wikitest:
27
  <table>
 
22
  It applies my RSF scale fitting technique to the Q_K quants and uses a different tensor mix layout compared to previously.
23
  6-June-2026: changed the MTP tensors to be NVFP4, ~130tk/s tg is now possible in some configurations on 5090.
24
  </p>
25
+ <p>I am experimenting with a smaller verison of this model designed to be used on machines with 16GB of VRAM:<BR>
26
+ <A HREF="https://huggingface.co/michaelw9999/Qwen3.6-27B-NVFP4-SMALL-MTP-GGUF">Qwen3.6-27B-NVFP4-SMALL-MTP-GGUF</A>.<BR>
27
+ The SMALL version will be slower than this model but should maintain similar quality, although I am still evaluating it.<BR></BR>
28
+
29
+ </p>
30
  <p>Feedback to improve this is appreciated.</p>
31
  Results below on RTX 5090 with Wikitest:
32
  <table>