Qwopus3.6-35B-A3B-v1-oQ4-mtp (oMLX build)

Local oMLX build of Jackrong/Qwopus3.6-35B-A3B-v1 — Qwen3.6 35B-A3B MoE (~3B active per token), text-only.

How it was built

  • Trunk: faithfully re-quantized from the bf16 source to oMLX oQ4 — 4-bit / group-size 64, with every MoE router gate (mlp.gate, mlp.shared_expert_gate) at 8-bit, matching the proven stamsam oQ4-MTP recipe. ~4.649 bits/weight.
  • MTP head: Qwopus's own MTP head is byte-identical to the unsloth base Qwen3.6-35B-A3B head, which drafts at 0% acceptance once quantized on this MoE (the known quantized-MTP collapse). This build therefore uses the distilled MTP head from Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled (stamsam), which is quantization-robust and achieves ~77% draft acceptance on this trunk.
  • The draft head only proposes tokens; the Qwopus trunk verifies and accepts/rejects each one, so generated content is 100% Qwopus — the head only affects speed.

Verified

Loads in oMLX as qwen3_5_moe (batched engine), native MTP patch (PR 990) active, MTP acceptance ~77% (60/78 measured). Trunk-only generation ~91 tok/s on M4 Max 48 GB.

Registered with mtp_enabled: true in ~/.omlx/model_settings.json. Build scripts: ~/.omlx-build/ (convert_qwopus.py, build_mtp.py, assemble_omlx.py, install_omlx.py).

Downloads last month
15
Safetensors
Model size
35B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support