michellemoorre commited on
Commit
65a71ce
·
verified ·
1 Parent(s): ad0031f

Refine checkpoint pipeline and card layout

Browse files
Files changed (2) hide show
  1. README.md +42 -11
  2. release-manifest.json +1 -1
README.md CHANGED
@@ -79,21 +79,52 @@ XS targets minimum footprint for non-thinking chat and instruction following. Fo
79
 
80
  </details>
81
 
82
- ## How we build the checkpoints
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
83
 
84
- Every tier starts from [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) at pinned revision [`851bf6e8`](https://huggingface.co/Qwen/Qwen3.5-4B/tree/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a) and passes through the same five-stage pipeline:
85
 
86
- | Stage | What happens |
87
- | --- | --- |
88
- | **1 · Pin the source** | Lock the upstream weights, tokenizer metadata, and model topology. |
89
- | **2 · Build the bank** | Produce native Q2_K–Q8_0 candidates for each quantizable tensor group and record their exact byte costs. |
90
- | **3 · Select and assemble** | Choose a model-specific XS/S assignment under an exact size cap; assemble M/L with fixed Q4_K/Q8_0 assignments. |
91
- | **4 · Tune scales** | Optimize representable scales and minima against teacher outputs while packed codes, dense weights, and tensor layout remain frozen. |
92
- | **5 · Validate the release** | Export GGUF, verify size and payload integrity, load with pinned llama.cpp, and run held-out KL before downstream benchmarks. |
93
 
94
- **Adaptive schedules for this model:** XS `RCO anchor`; S `RCO anchor`. `Uniform` scores module candidates independently; `anchor` uses activations from a concrete model trajectory. The S target matches the exact byte size of a pinned UD-Q2_K_XL checkpoint; its schedule and weights are produced by this pipeline.
95
 
96
- The recommended tier is selected independently for each base model from complete end-to-end evaluation—not from the tier name or nominal bit label.
97
 
98
  <details>
99
  <summary><b>Technical file details</b></summary>
 
79
 
80
  </details>
81
 
82
+ ## How the compression pipeline works
83
+
84
+ The adaptive tiers separate precision allocation from final weight reconstruction:
85
+
86
+ ```text
87
+ Q2_K … Q8_0 PTQ lanes
88
+
89
+ per-group candidate bank + native byte costs
90
+
91
+ ANNA-DIAG-BYTES or RCO allocation
92
+
93
+ fresh scheduled GPTQ / QuantEase + QEP
94
+
95
+ scale/min distillation with codes frozen
96
+ ```
97
+
98
+ ### 1. Build native candidates
99
+
100
+ For every quantizable tensor group, TorsionQuant produces six native K-quant candidates. Each lane uses the same Hessian-aware GPTQ stack: **QuantEase** optimizes the blockwise weight codes, **QEP** propagates reconstruction error through the layer, and native K refinement keeps the solver aligned with the GGUF layouts that will actually ship.
101
+
102
+ - A **uniform bank** runs each Q2/Q3/Q4/Q5/Q6/Q8 lane independently.
103
+ - An **anchor bank** builds every candidate from activations propagated along one concrete quantized trajectory.
104
+
105
+ ### 2. Allocate the byte budget
106
+
107
+ The scheduler sees the exact native GGUF cost of every `(tensor group, qtype)` choice, rather than a nominal average bit width.
108
+
109
+ - **ANNA-DIAG-BYTES** installs one candidate at a time in the BF16 model, measures full-vocabulary next-token KL against the teacher while every other group remains dense, and solves an exact multiple-choice knapsack over that isolated-loss table.
110
+ - **RCO** optimizes the choices jointly: projected-Gumbel search minimizes end-to-end teacher KL on the interpolated model under the byte constraint, then exact integer rounding produces the hard assignment.
111
+
112
+ | Adaptive tier | Bank / selector | Byte target |
113
+ | --- | --- | --- |
114
+ | XS | `RCO` · anchor bank | Compact model-specific cap |
115
+ | S | `RCO` · anchor bank | Exact size of a pinned `UD-Q2_K_XL` reference |
116
+
117
+ The S reference is used only as an external byte-budget authority.
118
+
119
+ ### 3. Requantize and refine
120
 
121
+ The bank determines the qtype map; the release weights are regenerated from the original dense model in a **fresh full-model PTQ pass**. Activation propagation and error compensation therefore follow the final mixed-precision trajectory.
122
 
123
+ A short affine distillation pass then tunes only the native scales and minima against cached teacher logits. Qtypes, packed integer codes, dense weights, and tensor layouts remain frozen, so the size and runtime contract cannot drift. M and L use the same reconstruction and refinement path with fixed Q4_K and Q8_0 assignments.
 
 
 
 
 
 
124
 
125
+ Only after these stages is the final GGUF exported and measured on the untouched 3,072-sequence KL set and the downstream benchmarks above. The recommendation is based on those end-to-end results, not on the scheduler's calibration objective.
126
 
127
+ **Base model:** [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) at revision [`851bf6e8`](https://huggingface.co/Qwen/Qwen3.5-4B/tree/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a).
128
 
129
  <details>
130
  <summary><b>Technical file details</b></summary>
release-manifest.json CHANGED
@@ -36,7 +36,7 @@
36
  },
37
  "display_name": "Qwen3.5 4B",
38
  "family": "Qwen 3.5",
39
- "generated_at": "2026-07-21T10:38:40.419392+00:00",
40
  "license": "apache-2.0",
41
  "model_key": "qwen3p5_4b",
42
  "reasoning_policy": {
 
36
  },
37
  "display_name": "Qwen3.5 4B",
38
  "family": "Qwen 3.5",
39
+ "generated_at": "2026-07-21T10:44:31.264168+00:00",
40
  "license": "apache-2.0",
41
  "model_key": "qwen3p5_4b",
42
  "reasoning_policy": {