How to use from the
Use from the
MLX library
# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm

# Generate text with mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("majentik/Qwen3.6-35B-A3B-RotorQuant-MLX-4bit")

prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True
)

text = generate(model, tokenizer, prompt=prompt, verbose=True)

KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively — use -ctk q8_0 -ctv q8_0 (half KV memory, negligible quality loss: perplexity +0.002–0.05) or -ctk q4_0 -ctv q4_0 (quarter memory, ≈7.6% perplexity increase). In Ollama: OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fast fused Flash-Attention path. Since April 2026, mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out: LLAMA_ATTN_ROT_DISABLE=1).

The RotorQuant/TurboQuant fork flow below is experimental/legacy: the TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork is unmaintained relative to mainline. It is NOT required to use this model.

Qwen3.6 35B-A3B - RotorQuant MLX 4-bit

4-bit weight-quantized MLX version of Qwen/Qwen3.6-35B-A3B with the legacy RotorQuant KV-cache fork (superseded by upstream llama.cpp KV options). Optimized for Apple Silicon inference via the MLX framework. A good balance between model quality and memory efficiency. Only 3B parameters are active per token despite 35B total, making this model significantly more efficient at inference time than its parameter count suggests.

Approximate model size: ~18 GB

Model Specifications

Property Value
Base Model Qwen/Qwen3.6-35B-A3B
Parameters 35 billion total (3 billion active per token)
Architecture Mixture-of-Experts (MoE) (3B active per token)
Modality Text-only (language tower extracted from a multimodal base; vision tower not included)
License Apache 2.0
Weight Quantization 4-bit (~18 GB)
KV-Cache Quantization RotorQuant
Framework MLX (Apple Silicon)

Quickstart

import mlx.core as mx
from mlx_lm import load, generate

model, tokenizer = load("majentik/Qwen3.6-35B-A3B-RotorQuant-MLX-4bit")

prompt = "Give me a short introduction to Mixture-of-Experts models."
response = generate(model, tokenizer, prompt=prompt, max_tokens=512)
print(response)

Text-only extraction. This repo contains only the quantized language tower of Qwen3.6-35B-A3B. The upstream vision tower (333 tensors) and MTP head are not included, so image/video input does not work and mlx_vlm.load(...) fails with a Missing ... parameters error (the vision tower it expects is absent from the checkpoint). Load it with mlx_lm (recent version with qwen3_5_moe support) as shown above. For image/video input, use the upstream BF16 model Qwen/Qwen3.6-35B-A3B on a runtime that supports it.

About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's release labels, not distinct quantization algorithms — for any given tier, both brand repos carry byte-identical weights produced with the standard MLX / llama.cpp quantizers. No brand-specific speedup is claimed or measured. The KV-cache fork these labels originally referred to is legacy; for KV-cache memory savings use the upstream options described above (-ctk/-ctv q8_0, OLLAMA_KV_CACHE_TYPE).

KV-Cache Quantization Comparison

Method Prefill Speed Decode Speed Memory Savings Reference
TurboQuant 1x (baseline) 1x (baseline) High arXiv: 2504.19874

Memory Estimates (Qwen3.6 35B-A3B)

Precision Approximate Size MLX Variant
FP16 (original) ~70 GB (approx.) --
8-bit quantized ~35 GB RotorQuant-MLX-8bit
4-bit quantized ~18 GB This model
2-bit quantized ~9 GB RotorQuant-MLX-2bit

Hardware Requirements

This model requires approximately 18 GB of unified memory. Recommended hardware:

  • Apple M2 Pro (24 GB+)
  • Apple M3 Pro (24 GB+)
  • Apple M4 Pro (24 GB+)
  • Any Apple Silicon Mac with 24 GB+ unified memory

See Also

Quant trade-off (MLX lane)

Bits Approx size Use case Recommendation
2-bit ~9.1 GB Aggressive quantization Very low-RAM Macs
3-bit ~13 GB Lossy but small Low-RAM Macs
4-bit ~15 GB Balanced default Recommended for most Macs
5-bit ~18 GB Higher fidelity Quality-sensitive
6-bit ~21 GB Approaching FP16 quality High-fidelity
8-bit ~27 GB Near-lossless reference Fidelity-critical work

(Current variant — 4bit — is bolded.)

Variants in this family

(Showing 24 sibling variants under majentik/qwen3.6-35b-a3b-*. The current variant — RotorQuant-MLX-4bit — is bolded.)

Variant Runtime Approx size Use case
RotorQuant-GGUF-IQ4_XS llama.cpp ~30 GB Lossy 4-bit, low-RAM CPU/edge
RotorQuant-GGUF-Q2_K llama.cpp ~21 GB Lossy, low-RAM CPU/edge
RotorQuant-GGUF-Q3_K_M llama.cpp ~27 GB Smaller 3-bit, CPU-friendly
RotorQuant-GGUF-Q4_K_M llama.cpp ~38 GB Balanced default
RotorQuant-GGUF-Q5_K_M llama.cpp ~46 GB Higher fidelity, more RAM
RotorQuant-GGUF-Q8_0 llama.cpp ~74 GB Near-lossless reference
RotorQuant-MLX-2bit mlx-lm ~11 GB Apple Silicon, smallest
RotorQuant-MLX-3bit mlx-lm ~16 GB Apple Silicon, small
RotorQuant-MLX-4bit mlx-lm ~22 GB Apple Silicon balanced
RotorQuant-MLX-5bit mlx-lm ~27 GB Apple Silicon, higher fidelity
RotorQuant-MLX-6bit mlx-lm ~32 GB Apple Silicon, near-lossless
RotorQuant-MLX-8bit mlx-lm ~41 GB Apple Silicon reference
TurboQuant-MLX-2bit mlx-lm ~11 GB Apple Silicon, smallest
TurboQuant-MLX-3bit mlx-lm ~16 GB Apple Silicon, small
TurboQuant-MLX-4bit mlx-lm ~22 GB Apple Silicon balanced
TurboQuant-MLX-5bit mlx-lm ~27 GB Apple Silicon, higher fidelity
TurboQuant-MLX-6bit mlx-lm ~32 GB Apple Silicon, near-lossless
TurboQuant-MLX-8bit mlx-lm ~41 GB Apple Silicon reference
Downloads last month
1,244
Safetensors
Model size
6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for majentik/Qwen3.6-35B-A3B-RotorQuant-MLX-4bit

Quantized
(669)
this model

Collection including majentik/Qwen3.6-35B-A3B-RotorQuant-MLX-4bit

Paper for majentik/Qwen3.6-35B-A3B-RotorQuant-MLX-4bit