Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-1M-qx86-hi-mlx

Glaciers

This model is a merge of:

  • armand0e/Qwen3.6-35B-A3B-Fable-5-Distill
  • Hcompany/Holo3.1-35B-A3B
  • Jackrong/Qwopus3.6-35B-A3B-Coder

Let me know if it worked for you.

If this model gets more Likes, I will provide the source--usually not here because of space constraints.

Brainwaves

         arc   arc/e boolq hswag obkqa piqa  wino
bf16     0.645,0.837,0.894,0.783,0.454,0.822,0.735
mxfp8    0.645,0.833,0.894,0.783,0.454,0.820,0.725
qx86-hi  0.647,0.843,0.893,0.780,0.446,0.822,0.730
qx64-hi  0.655,0.839,0.894,0.778,0.442,0.824,0.725
mxfp4    0.637,0.832,0.889,0.776,0.462,0.817,0.714

Similar model in this range

Jiunsong/SuperQwen-AgentWorld-35B-A3B-abliterated

         arc   arc/e boolq hswag obkqa piqa  wino
mxfp4    0.657,0.862,0.906,0.766,0.490,0.825,0.692

Qwen/Qwen-AgentWorld-35B-A3B

         arc   arc/e boolq hswag obkqa piqa  wino
qx64-hi  0.644,0.818,0.909
mxfp4    0.626,0.813,0.901

Model components

armand0e/Qwen3.6-35B-A3B-Fable-5-Distill

         arc   arc/e boolq hswag obkqa piqa  wino
qx86-hi  0.635,0.821,0.891,0.770,0.444,0.818,0.721

Hcompany/Holo-3.1-35B-A3B

         arc   arc/e boolq hswag obkqa piqa  wino
qx86-hi  0.533,0.705,0.882,0.771,0.456,0.811,0.690

Jackrong/Qwopus3.6-35B-A3B-Coder

         arc   arc/e boolq hswag obkqa piqa  wino
qx86-hi  0.594,0.770,0.888,0.750,0.438,0.813,0.717

Baseline model

Qwen3.6-35B-A3B-Instruct

         arc   arc/e boolq hswag obkqa piqa  wino
mxfp8    0.581,0.757,0.892,0.751,0.428,0.803,0.688
qx86-hi  0.576,0.742,0.896,0.745,0.422,0.803,0.708
mxfp4    0.586,0.767,0.886,0.751,0.428,0.798,0.681

Quant    Perplexity      Peak Memory   Tokens/sec
mxfp8    5.138 ± 0.037   42.65 GB      1201
mxfp4    5.158 ± 0.037   25.33 GB      1355
qx86-hi  4.826 ± 0.033   45.50 GB      1474
qx64-hi  4.710 ± 0.032   36.83 GB      1414

Thinking toggle

This model is using the fixed jinja template from froggeric/Qwen-Fixed-Chat-Templates

Contribute to NightmediaAI

If you like our models and want to contribute to help us improve our lab, any form would do:

ETH:0x6b6633606995BC180925c47d4249ED624aB7b2A5 USDC:0x19e6bDDCBa47BB09a9Bc153Bb6479fc57284421a BTC:36d7U1n3MFaXgnNRAaEL3Pa3Hy6oFhM7XY BCH:15dNMzhJ87XJSTU89VCBsDHj747QvBQaap

My models and I thank you :)

-G


Use with mlx

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-1M-qx86-hi-mlx")

prompt = "hello"

if tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_dict=False,
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)
Downloads last month
1,679
Safetensors
Model size
11B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-1M-qx86-hi-mlx