Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-dequantized-oQ4e-mtp

This model was quantized using oQ (oMLX v0.6.3) mixed-precision quantization.

Quantization details

  • Model type: qwen3_5_moe
  • Bits: 4
  • Group size: 64
  • Format: MLX safetensors

Performance Benchmark

  • Run on: Apple Mac Studio M4 Max 128GB
  • error
oMLX - LLM inference, optimized for your Mac
https://github.com/jundot/omlx
Benchmark Model: Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-dequantized-oQ4e-mtp
Engine: Auto
Context: Code (Python)
================================================================================

Single Request Results
--------------------------------------------------------------------------------
Test                                TTFT(ms)    TPOT(ms)        pp TPS        tg TPS      E2E(s)    Throughput    Peak Mem
pp1025/tg128                           577.6        8.67  1774.7 tok/s   116.3 tok/s       1.686   683.8 tok/s    20.02 GB
pp4097/tg128                          2211.3        9.08  1852.7 tok/s   111.0 tok/s       3.368  1254.4 tok/s    22.07 GB
pp8193/tg128                          4612.1        9.23  1776.4 tok/s   109.2 tok/s       5.792  1436.7 tok/s    22.70 GB
pp16385/tg128                        10047.0        9.96  1630.8 tok/s   101.2 tok/s      11.319  1458.9 tok/s    23.98 GB
pp32769/tg128                        23590.4       11.12  1389.1 tok/s    90.6 tok/s      25.011  1315.3 tok/s    26.53 GB
pp65537/tg128                        62547.3       13.50  1047.8 tok/s    74.7 tok/s      64.270  1021.7 tok/s    31.79 GB
pp131073/tg128                      190820.4       17.63   686.9 tok/s    57.2 tok/s     193.068   679.6 tok/s    42.25 GB

Continuous Batching
pp1024 / tg128
--------------------------------------------------------------------------------
Batch           tg TPS   Speedup        pp TPS    pp TPS/req    TTFT(ms)      E2E(s)
2x         204.8 tok/s       N/A  1201.2 tok/s   600.6 tok/s      1494.8       2.955
4x         331.5 tok/s       N/A  1084.0 tok/s   271.0 tok/s      2362.9       5.323
8x         564.3 tok/s       N/A  1012.8 tok/s   126.6 tok/s      4429.0       9.903

Intelligence Benchmark Comparison

Intelligence Benchmark Comparison

               Mode    Sampled                                                Qwen3.8-Flash-Next-oQ4e-mtp  Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-dequantized-oQ4e-mtp
-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------
MMLU           Sample  100/14042                                                                    90.0%                                                               87.5%
TRUTHFULQA     Full    817                                                                              -                                                               79.8%
HUMANEVAL      Full    164                                                                              -                                                               92.7%
LIVECODEBENCH  Sample  100/1055                                                                         -                                                               43.0%

--- Detail ---

Model: Qwen3.8-Flash-Next-oQ4e-mtp
Benchmark         Accuracy   Correct   Total   Time(s)   Think
--------------------------------------------------------------
MMLU                 90.0%        90     100    3288.9     Yes

Model: Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-dequantized-oQ4e-mtp
Benchmark         Accuracy   Correct   Total   Time(s)   Think
--------------------------------------------------------------
MMLU                 87.5%       875    1000    4108.8     Yes
TRUTHFULQA           79.8%       652     817    2650.3     Yes
HUMANEVAL            92.7%       152     164    1267.9     Yes
LIVECODEBENCH        43.0%        43     100    4244.5     Yes
Downloads last month
43
Safetensors
Model size
6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for symrex/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-dequantized-oQ4e-mtp