Flight Simulator — Coded by This Model

A real-world coding benchmark: each model was prompted to write a complete flight simulator from scratch. The resulting code was rendered and recorded.

Flight Simulator

Local SOTA for 48GB Macs — Intelligence Benchmark Comparison

This model is part of a benchmark comparison of the best local MLX-quantized LLMs that fit in 48GB unified memory on Apple Silicon. All benchmarks run in instruct mode (no thinking) with n=50 samples per benchmark.

SOTA Comparison

Benchmark Samples Agents-A1 6bit-XL Gemma-4 26B 6bit-XL Huihui-Qwen3.6 6bit-XL Ornith-35B 6bit-XL Qwen3.6-27B oQ4e Qwen3.6-35B 6bit-XL Qwen3.6-35B oQ4e Qwen3.6-35B oQ4e-XL Qwen3.6-35B oQ6
MMLU 50/14042 66% 76% 74% 64% 74% 64% 66% 72% 64%
MMLU_PRO 50/12032 58% 82% 66% 66% 56% 64% 60% 64% 60%
ARC_CHALLENGE 50/1172 90% 90% 92% 92% 88% 90% 92% 92% 90%
HUMANEVAL 50/164 90% 98% 84% 78% 92% 78% 92% 90% 66%
MBPP 50/500 70% 82% 78% 78% 86% 78% 80% 76% 76%
Average 74.8% 85.6% 78.8% 75.6% 79.2% 74.8% 78.0% 78.8% 71.2%

Collection: Local SOTA for 48GB Macs

⚠️ n=50 sampling means wide confidence intervals (±~13% at 95% CI). Differences under ~6 points may not be statistically significant. Models using data-aware quantization (oQ/oQe) may be calibrated on benchmark-like data — their scores carry a benchmaxxing caveat. The BaseQuant_XL variants (data-agnostic) provide the most honest generalization estimates.

leonsarmiento/Huihui-Qwen3.6-35B-A3B-abliterated-6bit-XL-mlx

This model was converted to MLX format from huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated using BaseQuant_XL 6/8-bit mixed quantization optimized for Apple Silicon. The vision encoder is preserved and quantized at 6-bit, making this a full multimodal model.

BaseQuant_XL keeps the most routing-critical layers in full bf16 precision — the MoE router gate, shared expert gate, shared expert, and lm_head — while applying aggressive quantization to the bulk parameters. This preserves routing accuracy and output quality where it matters most.

Qwen3.6-35B-A3B-abliterated is an uncensored version of Qwen3.6-35B-A3B where the refusal mechanism has been removed via directional ablation. It features 256 experts (8 active per token + 1 shared expert), hybrid full + linear (Gated DeltaNet) attention, a vision encoder, and an extended context window. Despite 35B total parameters, only ~3B are activated per token.

Use with mlx

pip install -U mlx-vlm
python -m mlx_vlm.generate --model leonsarmiento/Huihui-Qwen3.6-35B-A3B-abliterated-6bit-XL-mlx --max-tokens 256 --temperature 0.6 --top-p 0.95 --prompt "Hello"

BaseQuant_XL Quantization Strategy

Bit Depth Layers Rationale
bf16 (unquantized) mlp.gate (router), shared_expert_gate, lm_head, shared_expert Routing decisions and shared computation path — errors here are qualitatively different from precision loss
8-bit embed_tokens, self_attn (full attention), linear_attn (DeltaNet) Every-token layers with moderate sensitivity — 8-bit is near-lossless
6-bit vision_tower, switch_mlp (routed experts) Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision

Quantization Details

Layer Bits Group Size
mlp.gate (router) bf16
shared_expert_gate bf16
lm_head bf16
shared_expert bf16
embed_tokens 8 64
self_attn (full attention) 8 64
linear_attn (DeltaNet) 8 64
vision_tower 6 64
switch_mlp (routed experts) 6 64
Default fallback 8 64
  • Quantization type: BaseQuant_XL mixed (multimodal, vision preserved)
  • Bits per weight: 6.808
  • Total size: ~28 GB (6 shards)
  • Group size: 64
  • Method: Custom quant_predicate via mlx_vlm

Recommended Inference Parameters - Add to Jinja template on LM studio or Chat Template Kwargs on oMLX

Thinking Preserve ({%- set preserve_thinking = true %}):

  • General tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
  • Coding tasks: temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

Instruct Mode ({%- set enable_thinking = false -%}):

  • General tasks: temperature=0.7, top_p=0.8, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
  • Reasoning tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
Downloads last month
838
Safetensors
Model size
8B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for leonsarmiento/Huihui-Qwen3.6-35B-A3B-abliterated-6bit-XL-mlx

Quantized
(21)
this model

Collections including leonsarmiento/Huihui-Qwen3.6-35B-A3B-abliterated-6bit-XL-mlx