Qwen3.6-27B-MTP TQ3_4S — GGUF

TQ3_4S quant of Qwen/Qwen3.6-27B (Apache 2.0), produced via turbo-tan's llama.cpp-tq3 CUDA fork. Benchmarked on a Blackwell RTX 5060 Ti (16 GB). Mean AEON score: 0.525 (120 scored / 150 cases, 2048 budget).

File

File Size Quant BPW
Qwen3.6-27B-MTP-TQ3_4S.gguf 14.2 GB TQ3_4S (turbo-tan) ~3.8 bpw

NOT a stock llama.cpp quant

TQ3_4S (TurboQuant 3-bit 4-state) is a custom weight format unique to turbo-tan/llama.cpp-tq3. Stock llama.cpp and nixpkgs llama-cpp will exit with unknown quantization at load time. Use the llama-server/llama-cli from the tq3 fork.

Scope of these benchmarks — read this first

These numbers are a light baseline, not a thorough TQ3 evaluation. The mesh's bench framework is built for production agent workload regression-detection on the local stack, not for the kind of multi-axis sweep that upstream quant maintainers typically publish. Specifically:

  • Harness scope is bounded. The numbers below come from the mesh's Aeon-Bench-Pod (self-reported mode, 150-case full suite, text-only, no agentic harness). That's a regression suite, not a quality benchmark.
  • Sample sizes are small. 150 cases on a single GPU, single rep. None are powered for statistical significance.
  • No perplexity / wikitext / MMLU / GSM8K. The mesh's stack isn't a quality benchmark — those are upstream's territory.
  • Single GPU class (Blackwell 16 GB). All measurements are on an NVIDIA RTX 5060 Ti 16 GB (CUDA 13.2, turbo-tan/llama.cpp-tq3). No RDNA4, no multi-GPU, no Vulkan. Cross-hardware generalization is NOT implied.
  • No human eval. "0.525 mean on the AEON suite" is not a quality verdict on this specific quant.
  • 30 prose cases unscored — no frontier judge endpoint configured, so the prose category (30/150 cases) is excluded from the mean.

What this IS good for: a quick signal that the quant (a) loads, (b) runs at sane throughput, (c) produces coherent output across math, instruction, reasoning, and coding. What this is NOT good for: claiming "this is the best quant of this model," reproducing academic benchmark results, or substituting for upstream's validation work.

For a rigorous view, see Qwen/Qwen3.6-27B (parent model) and turbo-tan/llama.cpp-tq3 (quantizer). Raw bench reports are attached as BENCH-*.md files in this repo.

What we measured

AEON v3 — 150 cases (120 scored)

Benchmarked on Blackwell RTX 5060 Ti 16 GB, full suite.

Category Mean N
Overall 0.525 120
coding 0.933 30
math 0.467 30
reasoning 0.367 30
instruction 0.333 30
Difficulty Mean N
easy 1.000 8
medium 0.833 12
hard 0.700 20
expert 0.500 32
frontier 0.312 48

Profile: Strong on coding (0.933), weak on instruction-following (0.333) and reasoning (0.367). Typical Qwen3.6-TQ3 profile — excels at structured code tasks, struggles on instruction-heavy and frontier-difficulty cases.

Companion: ROCmFP4_FAST (RDNA4)

The companion ROCmFP4_FAST quant for AMD RDNA4 scores 0.558 on the same suite. Cross-format parity is within ~3pp on most categories; the AMD arm benefits from a higher coding score but the overall profile is similar.

Quick start

# Build turbo-tan's tq3 fork
git clone https://github.com/turbo-tan/llama.cpp-tq3
cd llama.cpp-tq3
mkdir build && cd build
cmake .. -DGGML_CUDA=ON -DCUDA_DOCKER_ARCH=sm_120
make -j$(nproc)

# Serve with production-equivalent flags
llama-server \
  -m Qwen3.6-27B-MTP-TQ3_4S.gguf \
  --host 0.0.0.0 --port 8081 \
  -ngl 99 -c 131072 -t 12 \
  -ctk q4_0 -ctv q4_0 \
  -np 1 --batch-size 512 --ubatch-size 128 \
  --jinja --metrics -rea off

Reproduce the quant

# Requires the tq3 fork and the BF16 source GGUF
llama-quantize --allow-requantize Qwen3.6-27B-BF16.gguf \
  Qwen3.6-27B-MTP-TQ3_4S.gguf TQ3_4S

Files in this repo

File Description
Qwen3.6-27B-MTP-TQ3_4S.gguf The quantized model (LFS-tracked)
README.md This model card
BENCH-aeon-full-suite.md AEON bench results (150 cases, 120 scored)

What's NOT in this repo (caveats)

  • Stock llama.cpp will not load this file. TQ3_4S is a custom weight format unique to turbo-tan/llama.cpp-tq3.
  • No AMD GPU bench. All measurements are RTX 5060 Ti (CUDA). The companion ROCmFP4_FAST quant for RDNA4 is in a separate repo.
  • No quality benchmark (perplexity, MMLU, GSM8K). The custom 3-bit quant works on the mesh's regression tests; whether it's "the best TQ3 quant" needs upstream validation.
  • No MTP / speculative-decode bench. MTP heads are present in the source model but were not benched on this quant.

Provenance

  • Source model: Qwen/Qwen3.6-27B — Apache 2.0
  • Quantizer: turbo-tan/llama.cpp-tq3
  • Quantizer license: MIT
  • Build hardware: NVIDIA RTX 5060 Ti 16 GB (Blackwell), CUDA 13.2, NixOS 25.11
  • Bench harness: Aeon-Bench-Pod v1 (self-reported mode)

License

  • Qwen/Qwen3.6-27B is Apache 2.0.
  • turbo-tan/llama.cpp-tq3 is MIT.
  • The GGUF in this repo is a derivative of the Apache 2.0-licensed parent, produced with the MIT-licensed quantizer. The Apache 2.0 license is preserved.
Downloads last month
5,495
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for maczzinatui/Qwen3.6-27B-MTP-TQ3_4S-GGUF

Base model

Qwen/Qwen3.6-27B
Quantized
(657)
this model

Collection including maczzinatui/Qwen3.6-27B-MTP-TQ3_4S-GGUF