Qwen3.5-4B-GGUF / README.md
michellemoorre's picture
Polish model card narrative and citations
cdb6df7 verified
|
Raw
History Blame
10.4 kB
metadata
license: apache-2.0
base_model:
  - Qwen/Qwen3.5-4B
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
thumbnail: >-
  https://huggingface.co/TheStageAI/Qwen3.5-4B-GGUF/resolve/main/assets/thestage-edge-models-header.png
tags:
  - gguf
  - llama.cpp
  - quantized
  - mixed-precision
  - local-inference
  - qwen3.5

TheStageAI Edge Models — the right model at every memory budget

Qwen3.5 4B — TheStageAI GGUF

Four deployment tiers for local inference with llama.cpp · 1.52 GB–4.49 GB
Start with M: 2.39 GB and 98% of BF16 instruction-strict IFEval.

Explore edge-lm on GitHub Read TheStageAI documentation Open TheStageAI Platform

Qwen 3.5 family: 0.8B  ·  2B  ·  4B  ·  9B

Start here

Tier Size Best for File
XS 1.52 GB Minimum footprint Download
S 1.90 GB Compact Download
M 2.39 GB Recommended · Balanced Download
L 4.49 GB High-precision Q8 Download

Exact byte counts, SHA-256 hashes, and tensor metadata: release-manifest.json.

Run with llama.cpp

llama-cli \
  --hf-repo TheStageAI/Qwen3.5-4B-GGUF \
  --hf-file Qwen3.5-4B-M-TS-Q4_K_M.gguf

Why M is the default

At 2.39 GB, M retains 98.4% of the BF16 instruction-strict IFEval score; on MMLU-Pro, it retains 99.1% of the BF16 score. It uses 47% less disk than L, making it the default for this release.

Tier IFEval strict — prompt / instruction (%) MMLU-Pro (%)
BF16 reference 82.44 / 87.53 79.55
XS 70.43 / 78.30
S 77.82 / 83.93 74.39
M 80.22 / 86.09 78.86
L 81.70 / 87.05 79.59

IFEval measures deterministic non-thinking instruction following; MMLU-Pro measures sampled long-form reasoning. Only complete scores are shown; means not reported.

XS and reasoning: use S, M, or L for long-form reasoning. Run Qwen XS with --reasoning off.

Evaluation protocol
  • IFEval: 541 prompts, native chat template, enable_thinking=false, temperature 0.
  • MMLU-Pro: 12,032 questions, native chat template, enable_thinking=true, temperature 1, top-p 0.95, 32,768-token output limit.
  • The headline BF16 comparison uses instruction-strict IFEval; the raw scores are shown in the table.

In matched long-thinking diagnostics, XS produced longer trajectories and reached the 32,768-token output limit more often than S. The XS MMLU-Pro cell is marked for that reason.

From source weights to deployment tiers

All four tiers are produced by the same production PTQ pipeline; only the precision map changes. The process moves from native code fitting, through sequential reconstruction and budget-aware scheduling, to a final model-wide alignment pass.

1. Fit native discrete codes

Calibration activations define a curvature objective weighted by true Fisher information for each quantized projection. NeUQI initializes every affine group's scale and minimum on that objective, so sensitive weight directions influence the grid more strongly.

With the grid fixed, a guarded cyclic coordinate-descent solver inspired by QuantEase searches the integer codes. Continuous sweeps can escape a poor initial projection; projected sweeps return to a valid discrete solution. Round-to-nearest remains a non-regression baseline, and a final K-quant refinement optimizes the stored scales and minima while keeping packed codes fixed.

2. Reconstruct the trajectory the model will run

Layers are processed in execution order. Every projection is calibrated against activations from the already-quantized prefix, while a dense reference path measures accumulated drift. Quantization Error Propagation (QEP) folds that drift into the next reconstruction target, allowing later layers to compensate for errors they will actually receive at inference time.

3. Allocate the encoded byte budget

For XS and S, each quantizable group can select among native Q2_K through Q8_0 representations. The schedule optimizer trades changes in the teacher distribution against exact encoded byte cost, including scale and minimum metadata. Sensitive groups keep more precision; robust groups carry more compression.

ANNA provides TheStageAI's automated constrained configuration search, while RCO supplies an exact-budget search route (reference implementation). For this model, RCO selected both the XS and S precision maps. M and L use fixed Q4_K_M and Q8_0 decoder qtype choices, respectively, while retaining the same reconstruction and scale-tuning stages.

Once the map is selected, PTQ is rerun from the original source weights. Every layer therefore sees the final upstream precision choices rather than a collection of independently prepared bank tensors.

4. Align the complete model

A short affine distillation pass freezes qtypes, packed codes, dense weights, and tensor layouts while tuning native FP16 scales and minima. The loss matches the teacher's next-token distribution—including high-probability tokens and the remaining tail mass—without changing file size or runtime layout.

Every shipping GGUF is hashed, load-tested, and evaluated on a held-out set of 3,072 sequences using next-token KL. Downstream harnesses use a deterministic HF mirror reconstructed from that exact GGUF; its source SHA-256 and evaluation IDs are recorded in release-manifest.json. The recommended tier is chosen from complete-model results, not from a local reconstruction proxy.

Technical file details
Tier Hub selector GGUF file type Whole-file BPW
XS Q3_K_S MOSTLY_Q2_K 2.889
S Q4_K_S MOSTLY_Q2_K 3.622
M Q4_K_M MOSTLY_Q4_K_M 4.538
L Q8_0 MOSTLY_Q8_0 8.533

The Hub selector controls sidebar grouping and download discovery. For XS and S it approximates the whole-file size class; release-manifest.json is authoritative for the internal tensor mix. M and L use fixed Q4_K_M and Q8_0 decoder qtype choices within the same production PTQ pipeline.

Runtime memory also includes KV cache and buffers, which grow with context length.

TheStageAI Edge Stack

  • Portable local inference: these GGUF files for llama.cpp-compatible runtimes.
  • Native Apple Silicon: edge-lm for compressed MLX models on Macs and iPhones.
  • Automated compression search: ANNA for budget-constrained configuration discovery.
  • Custom deployment: the TheStageAI Platform and documentation for compression, compilation, and serving workflows.

Have a device, latency, or memory target? Talk to our team →

Reproducibility

Citation

If you use this checkpoint, please cite the upstream base model and this release:

@misc{thestageai2026qwen3p54bgguf,
  author       = {{TheStageAI}},
  title        = {Qwen3.5 4B — TheStageAI GGUF Release},
  year         = {2026},
  month        = {jul},
  howpublished = {Hugging Face model release},
  url          = {https://huggingface.co/TheStageAI/Qwen3.5-4B-GGUF},
  note         = {XS, S, M, and L deployment tiers},
}

Methods and tools

License

The model weights are released under the upstream model's Apache-2.0 license. llama.cpp and other runtime software retain their own licenses.