michellemoorre commited on
Commit
cdb6df7
·
verified ·
1 Parent(s): 4d5f4ee

Polish model card narrative and citations

Browse files
Files changed (2) hide show
  1. README.md +55 -25
  2. release-manifest.json +3 -3
README.md CHANGED
@@ -22,9 +22,9 @@ tags:
22
  <h1 align="center">Qwen3.5 4B — TheStageAI GGUF</h1>
23
 
24
  <p align="center">
25
- Four GGUF deployment points for llama.cpp · 1.52 GB–4.49 GB
26
  <br>
27
- <strong>Recommended: M · 2.39 GB · 98% of BF16 IFEval</strong>
28
  </p>
29
 
30
  <div style="display: flex; gap: 8px; justify-content: center; align-items: center; margin: 16px 0;">
@@ -42,7 +42,7 @@ tags:
42
  | XS | 1.52 GB | Minimum footprint | [Download](./Qwen3.5-4B-XS-TS-Q3_K_S.gguf) |
43
  | S | 1.90 GB | Compact | [Download](./Qwen3.5-4B-S-TS-Q4_K_S.gguf) |
44
  | **M** | **2.39 GB** | **Recommended · Balanced** | [Download](./Qwen3.5-4B-M-TS-Q4_K_M.gguf) |
45
- | L | 4.49 GB | Maximum fidelity | [Download](./Qwen3.5-4B-L-TS-Q8_0.gguf) |
46
 
47
  Exact byte counts, SHA-256 hashes, and tensor metadata: [`release-manifest.json`](./release-manifest.json).
48
 
@@ -54,7 +54,9 @@ llama-cli \
54
  --hf-file Qwen3.5-4B-M-TS-Q4_K_M.gguf
55
  ```
56
 
57
- ## Quality
 
 
58
 
59
  | Tier | IFEval strict — prompt / instruction (%) | MMLU-Pro (%) |
60
  | --- | ---: | ---: |
@@ -64,7 +66,7 @@ llama-cli \
64
  | **M** | 80.22 / 86.09 | 78.86 |
65
  | L | 81.70 / 87.05 | 79.59 |
66
 
67
- **Modes.** IFEval uses 541 deterministic non-thinking prompts. MMLU-Pro covers 12,032 questions with sampled long reasoning. Only complete MMLU-Pro runs are shown; `—` means not reported.
68
 
69
  > **XS and reasoning:** use S, M, or L for long-form reasoning. Run Qwen XS with `--reasoning off`.
70
 
@@ -73,67 +75,95 @@ llama-cli \
73
 
74
  - **IFEval:** 541 prompts, native chat template, `enable_thinking=false`, temperature 0.
75
  - **MMLU-Pro:** 12,032 questions, native chat template, `enable_thinking=true`, temperature 1, top-p 0.95, 32,768-token output limit.
76
- - The headline BF16-retention percentage uses instruction-strict IFEval and is capped at 100%; the uncapped raw scores are shown in the table.
77
 
78
- XS targets minimum footprint for non-thinking chat and instruction following. For long-form reasoning, use S, M, or L. Run Qwen XS with `--reasoning off`. In matched long-thinking diagnostics, XS produced longer trajectories and reached the 32,768-token output limit more often than S. The XS MMLU-Pro cell is consequently marked `—`.
79
 
80
  </details>
81
 
82
- ## Inside the production PTQ
83
 
84
- At 2–4 bits, the main challenge is not storing fewer values. It is keeping a small local reconstruction error from changing the activation distribution of every layer that follows. The production pipeline handles that problem at four levels: the quantization grid, the layer trajectory, the model-wide precision map, and the final output distribution.
85
 
86
- ### 1. Fit discrete codes on the native grid
87
 
88
- Calibration activations are used to build a true-Fisher-weighted curvature matrix for every quantized projection. **NeUQI** initializes each affine group range from that curvature, giving important weight directions more influence over the scale and minimum.
89
 
90
- With the grid fixed, a guarded **QuantEase-like cyclic coordinate-descent solver** searches the integer weight codes. Each coordinate is updated against the current choices for all other coordinates; continuous relaxation sweeps help escape a poor early projection, while projected sweeps return a valid discrete solution. The best feasible assignment is kept with round-to-nearest as a non-regression baseline. A final native K refinement improves the stored scale/min side parameters against the same curvature objective, with the packed codes held fixed.
91
 
92
- ### 2. Reconstruct along the quantized trajectory
93
 
94
- The model is processed in execution order, not as a collection of independent matrices. Each later projection receives activations from the already quantized prefix. In parallel, the dense reference path measures the activation drift accumulated so far; **quantization error propagation (QEP)** folds that drift into the next reconstruction target. Later layers can therefore absorb part of the upstream error instead of being optimized for activations they will never see at inference time.
95
 
96
- ### 3. Spend bytes where they preserve the model
97
 
98
- XS and S expose native Q2_K through Q8_0 choices for every quantizable group. Their mixed-precision map is selected from the effect of those choices on the teacher distribution under the **exact encoded byte cost**, including scale/min side data. Sensitive groups retain more precision, while robust groups carry more of the compression. M and L keep fixed Q4_K and Q8_0 maps as the higher-precision operating points.
99
 
100
- After the map is chosen, the release is not assembled from individually quantized bank tensors. The sequential PTQ pass is rerun from the original source weights so every reconstruction sees the final upstream precision choices.
101
 
102
- ### 4. Align the complete model without changing its format
103
 
104
- A short affine distillation pass keeps qtypes, packed weight and side codes, dense weights, and tensor layouts fixed while tuning only the native FP16 scales and minima. Its loss matches the teacher's next-token distribution—including the high-probability tokens and the remaining tail mass—so the final correction is model-wide while file size and runtime layout stay unchanged.
105
 
106
- The exported GGUF is then load-tested and evaluated as the artifact that ships: first on an untouched 3,072-sequence next-token KL set, then on the downstream benchmarks above. The recommended tier is chosen from those complete-model results, not from a local reconstruction proxy.
107
 
108
- **Base model:** [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) at revision [`851bf6e8`](https://huggingface.co/Qwen/Qwen3.5-4B/tree/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a).
109
 
110
  <details>
111
  <summary><b>Technical file details</b></summary>
112
 
113
- | Tier | Hub label | GGUF file type | Whole-file BPW |
114
  | --- | --- | --- | ---: |
115
  | XS | `Q3_K_S` | `MOSTLY_Q2_K` | 2.889 |
116
  | S | `Q4_K_S` | `MOSTLY_Q2_K` | 3.622 |
117
  | M | `Q4_K_M` | `MOSTLY_Q4_K_M` | 4.538 |
118
  | L | `Q8_0` | `MOSTLY_Q8_0` | 8.533 |
119
 
120
- For XS and S, the public Q-type suffix is an approximate Hub size label for discoverability; [`release-manifest.json`](./release-manifest.json) is authoritative for tensor types. M and L use fixed Q4_K and Q8_0 assignments for quantized decoder tensors.
121
 
122
  Runtime memory also includes KV cache and buffers, which grow with context length.
123
 
124
  </details>
125
 
126
- ## More from TheStageAI
127
 
128
- These GGUFs provide a portable llama.cpp deployment path. For native compressed models on Apple Silicon, explore [edge-lm](https://github.com/TheStageAI/edge-lm). For custom compression, compilation, and deployment workflows, use the [TheStageAI Platform](https://app.thestage.ai/) and [documentation](https://docs.thestage.ai/).
 
 
 
129
 
130
  **Have a device, latency, or memory target? [Talk to our team →](https://app.thestage.ai/contact)**
131
 
132
  ## Reproducibility
133
 
 
 
134
  - **Manifest:** [`release-manifest.json`](./release-manifest.json) records the exact base revision, byte sizes, GGUF file types, whole-file BPW, tensor inventories, SHA-256 digests, held-out KL values, and evaluation IDs.
135
  - **Runtime gate:** export and load checks used [llama.cpp revision `bec4772f`](https://github.com/ggml-org/llama.cpp/commit/bec4772f6a2527d371557b5d2032641e5ff7619c).
136
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
137
  ## License
138
 
139
  The model weights are released under the upstream model's **Apache-2.0** license. llama.cpp and other runtime software retain their own licenses.
 
22
  <h1 align="center">Qwen3.5 4B — TheStageAI GGUF</h1>
23
 
24
  <p align="center">
25
+ Four deployment tiers for local inference with llama.cpp · 1.52 GB–4.49 GB
26
  <br>
27
+ <strong>Start with M: 2.39 GB and 98% of BF16 instruction-strict IFEval.</strong>
28
  </p>
29
 
30
  <div style="display: flex; gap: 8px; justify-content: center; align-items: center; margin: 16px 0;">
 
42
  | XS | 1.52 GB | Minimum footprint | [Download](./Qwen3.5-4B-XS-TS-Q3_K_S.gguf) |
43
  | S | 1.90 GB | Compact | [Download](./Qwen3.5-4B-S-TS-Q4_K_S.gguf) |
44
  | **M** | **2.39 GB** | **Recommended · Balanced** | [Download](./Qwen3.5-4B-M-TS-Q4_K_M.gguf) |
45
+ | L | 4.49 GB | High-precision Q8 | [Download](./Qwen3.5-4B-L-TS-Q8_0.gguf) |
46
 
47
  Exact byte counts, SHA-256 hashes, and tensor metadata: [`release-manifest.json`](./release-manifest.json).
48
 
 
54
  --hf-file Qwen3.5-4B-M-TS-Q4_K_M.gguf
55
  ```
56
 
57
+ ## Why M is the default
58
+
59
+ At 2.39 GB, M retains 98.4% of the BF16 instruction-strict IFEval score; on MMLU-Pro, it retains 99.1% of the BF16 score. It uses 47% less disk than L, making it the default for this release.
60
 
61
  | Tier | IFEval strict — prompt / instruction (%) | MMLU-Pro (%) |
62
  | --- | ---: | ---: |
 
66
  | **M** | 80.22 / 86.09 | 78.86 |
67
  | L | 81.70 / 87.05 | 79.59 |
68
 
69
+ IFEval measures deterministic non-thinking instruction following; MMLU-Pro measures sampled long-form reasoning. Only complete scores are shown; `—` means not reported.
70
 
71
  > **XS and reasoning:** use S, M, or L for long-form reasoning. Run Qwen XS with `--reasoning off`.
72
 
 
75
 
76
  - **IFEval:** 541 prompts, native chat template, `enable_thinking=false`, temperature 0.
77
  - **MMLU-Pro:** 12,032 questions, native chat template, `enable_thinking=true`, temperature 1, top-p 0.95, 32,768-token output limit.
78
+ - The headline BF16 comparison uses instruction-strict IFEval; the raw scores are shown in the table.
79
 
80
+ In matched long-thinking diagnostics, XS produced longer trajectories and reached the 32,768-token output limit more often than S. The XS MMLU-Pro cell is marked `—` for that reason.
81
 
82
  </details>
83
 
84
+ ## From source weights to deployment tiers
85
 
86
+ All four tiers are produced by the same production PTQ pipeline; only the precision map changes. The process moves from **native code fitting**, through **sequential reconstruction** and **budget-aware scheduling**, to a final **model-wide alignment** pass.
87
 
88
+ ### 1. Fit native discrete codes
89
 
90
+ Calibration activations define a curvature objective weighted by true Fisher information for each quantized projection. [NeUQI](https://arxiv.org/abs/2505.17595) initializes every affine group's scale and minimum on that objective, so sensitive weight directions influence the grid more strongly.
91
 
92
+ With the grid fixed, a guarded cyclic coordinate-descent solver inspired by [QuantEase](https://arxiv.org/abs/2309.01885) searches the integer codes. Continuous sweeps can escape a poor initial projection; projected sweeps return to a valid discrete solution. Round-to-nearest remains a non-regression baseline, and a final K-quant refinement optimizes the stored scales and minima while keeping packed codes fixed.
93
 
94
+ ### 2. Reconstruct the trajectory the model will run
95
 
96
+ Layers are processed in execution order. Every projection is calibrated against activations from the already-quantized prefix, while a dense reference path measures accumulated drift. [Quantization Error Propagation (QEP)](https://arxiv.org/abs/2504.09629) folds that drift into the next reconstruction target, allowing later layers to compensate for errors they will actually receive at inference time.
97
 
98
+ ### 3. Allocate the encoded byte budget
99
 
100
+ For XS and S, each quantizable group can select among native Q2_K through Q8_0 representations. The schedule optimizer trades changes in the teacher distribution against **exact encoded byte cost**, including scale and minimum metadata. Sensitive groups keep more precision; robust groups carry more compression.
101
 
102
+ [ANNA](https://docs.thestage.ai/qlip/docs/source/anna_api.html) provides TheStageAI's automated constrained configuration search, while [RCO](https://arxiv.org/abs/2605.00649) supplies an exact-budget search route ([reference implementation](https://github.com/IST-DASLab/RCO)). For this model, [RCO](https://arxiv.org/abs/2605.00649) selected both the XS and S precision maps. M and L use fixed Q4_K_M and Q8_0 decoder qtype choices, respectively, while retaining the same reconstruction and scale-tuning stages.
103
 
104
+ Once the map is selected, PTQ is rerun from the original source weights. Every layer therefore sees the final upstream precision choices rather than a collection of independently prepared bank tensors.
105
 
106
+ ### 4. Align the complete model
107
 
108
+ A short affine distillation pass freezes qtypes, packed codes, dense weights, and tensor layouts while tuning native FP16 scales and minima. The loss matches the teacher's next-token distribution—including high-probability tokens and the remaining tail mass—without changing file size or runtime layout.
109
 
110
+ Every shipping GGUF is hashed, load-tested, and evaluated on a held-out set of 3,072 sequences using next-token KL. Downstream harnesses use a deterministic HF mirror reconstructed from that exact GGUF; its source SHA-256 and evaluation IDs are recorded in [`release-manifest.json`](./release-manifest.json). The recommended tier is chosen from complete-model results, not from a local reconstruction proxy.
111
 
112
  <details>
113
  <summary><b>Technical file details</b></summary>
114
 
115
+ | Tier | Hub selector | GGUF file type | Whole-file BPW |
116
  | --- | --- | --- | ---: |
117
  | XS | `Q3_K_S` | `MOSTLY_Q2_K` | 2.889 |
118
  | S | `Q4_K_S` | `MOSTLY_Q2_K` | 3.622 |
119
  | M | `Q4_K_M` | `MOSTLY_Q4_K_M` | 4.538 |
120
  | L | `Q8_0` | `MOSTLY_Q8_0` | 8.533 |
121
 
122
+ The Hub selector controls sidebar grouping and download discovery. For XS and S it approximates the whole-file size class; [`release-manifest.json`](./release-manifest.json) is authoritative for the internal tensor mix. M and L use fixed Q4_K_M and Q8_0 decoder qtype choices within the same production PTQ pipeline.
123
 
124
  Runtime memory also includes KV cache and buffers, which grow with context length.
125
 
126
  </details>
127
 
128
+ ## TheStageAI Edge Stack
129
 
130
+ - **Portable local inference:** these GGUF files for llama.cpp-compatible runtimes.
131
+ - **Native Apple Silicon:** [edge-lm](https://github.com/TheStageAI/edge-lm) for compressed MLX models on Macs and iPhones.
132
+ - **Automated compression search:** [ANNA](https://docs.thestage.ai/qlip/docs/source/anna_api.html) for budget-constrained configuration discovery.
133
+ - **Custom deployment:** the [TheStageAI Platform](https://app.thestage.ai/) and [documentation](https://docs.thestage.ai/) for compression, compilation, and serving workflows.
134
 
135
  **Have a device, latency, or memory target? [Talk to our team →](https://app.thestage.ai/contact)**
136
 
137
  ## Reproducibility
138
 
139
+ - **Release:** July 21, 2026.
140
+ - **Base model:** [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) at revision [`851bf6e8`](https://huggingface.co/Qwen/Qwen3.5-4B/tree/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a).
141
  - **Manifest:** [`release-manifest.json`](./release-manifest.json) records the exact base revision, byte sizes, GGUF file types, whole-file BPW, tensor inventories, SHA-256 digests, held-out KL values, and evaluation IDs.
142
  - **Runtime gate:** export and load checks used [llama.cpp revision `bec4772f`](https://github.com/ggml-org/llama.cpp/commit/bec4772f6a2527d371557b5d2032641e5ff7619c).
143
 
144
+ ## Citation
145
+
146
+ If you use this checkpoint, please cite the upstream base model and this release:
147
+
148
+ ```bibtex
149
+ @misc{thestageai2026qwen3p54bgguf,
150
+ author = {{TheStageAI}},
151
+ title = {Qwen3.5 4B — TheStageAI GGUF Release},
152
+ year = {2026},
153
+ month = {jul},
154
+ howpublished = {Hugging Face model release},
155
+ url = {https://huggingface.co/TheStageAI/Qwen3.5-4B-GGUF},
156
+ note = {XS, S, M, and L deployment tiers},
157
+ }
158
+ ```
159
+
160
+ ### Methods and tools
161
+
162
+ - **Schedule selection:** [ANNA](https://docs.thestage.ai/qlip/docs/source/anna_api.html), TheStageAI's automated constrained compression configuration search.
163
+ - **Exact-budget optimization:** [RCO: Model Compression with Exact Budget Constraints via Riemannian Manifolds](https://arxiv.org/abs/2605.00649) ([code](https://github.com/IST-DASLab/RCO)).
164
+ - **Discrete PTQ:** [NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs](https://arxiv.org/abs/2505.17595) and [QuantEase: Optimization-based Quantization for Language Models](https://arxiv.org/abs/2309.01885).
165
+ - **Sequential reconstruction:** [Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization](https://arxiv.org/abs/2504.09629).
166
+
167
  ## License
168
 
169
  The model weights are released under the upstream model's **Apache-2.0** license. llama.cpp and other runtime software retain their own licenses.
release-manifest.json CHANGED
@@ -36,7 +36,7 @@
36
  },
37
  "display_name": "Qwen3.5 4B",
38
  "family": "Qwen 3.5",
39
- "generated_at": "2026-07-21T10:55:03.132704+00:00",
40
  "license": "apache-2.0",
41
  "model_key": "qwen3p5_4b",
42
  "reasoning_policy": {
@@ -185,7 +185,7 @@
185
  "product": "M",
186
  "reasoning_support": "supported",
187
  "recommended": true,
188
- "schedule_method": "fixed Q4_K",
189
  "sha256": "f8e45572b9cc35161d4772b09bccfd383fe0bb03fc6d69b40a9138731302290b",
190
  "tensor_count": 426,
191
  "tensor_type_counts": {
@@ -225,7 +225,7 @@
225
  "lm_head_policy": "skip_tied",
226
  "output_embedding_mode": "tied_alias",
227
  "parameter_count": 4205751296,
228
- "positioning": "Maximum fidelity",
229
  "product": "L",
230
  "reasoning_support": "supported",
231
  "recommended": false,
 
36
  },
37
  "display_name": "Qwen3.5 4B",
38
  "family": "Qwen 3.5",
39
+ "generated_at": "2026-07-21T12:22:14.106699+00:00",
40
  "license": "apache-2.0",
41
  "model_key": "qwen3p5_4b",
42
  "reasoning_policy": {
 
185
  "product": "M",
186
  "reasoning_support": "supported",
187
  "recommended": true,
188
+ "schedule_method": "fixed Q4_K_M",
189
  "sha256": "f8e45572b9cc35161d4772b09bccfd383fe0bb03fc6d69b40a9138731302290b",
190
  "tensor_count": 426,
191
  "tensor_type_counts": {
 
225
  "lm_head_policy": "skip_tied",
226
  "output_embedding_mode": "tied_alias",
227
  "parameter_count": 4205751296,
228
+ "positioning": "High-precision Q8",
229
  "product": "L",
230
  "reasoning_support": "supported",
231
  "recommended": false,