File size: 16,305 Bytes
31596cd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
# Reproducing this GGUF release

The release is pinned to immutable source, converter, calibration, and
validation revisions. BailingMoE3 support is not yet merged in upstream
`llama.cpp`; use the exact revision below.

## 1. Build the converter and runtime

```bash
git clone https://github.com/aetherbird/llama.cpp.git
git -C llama.cpp checkout d8d862521e9ad842f2b47f3b392b039317782aa0

python3 -m venv .venv
. .venv/bin/activate
python -m pip install -r llama.cpp/requirements.txt
python -m pip install --extra-index-url https://download.pytorch.org/whl/cu128 \
  gguf==0.19.0 huggingface_hub==0.36.2 numpy==1.26.4 \
  protobuf==4.25.9 safetensors==0.8.0 sentencepiece==0.2.2 \
  tokenizers==0.22.2 'torch==2.8.0+cu128' tqdm==4.70.0 \
  transformers==4.57.6

cmake -S llama.cpp -B llama.cpp/build \
  -DCMAKE_BUILD_TYPE=Release \
  -DGGML_CUDA=ON \
  -DGGML_NATIVE=OFF \
  -DCMAKE_CUDA_ARCHITECTURES=native
cmake --build llama.cpp/build --config Release --parallel \
  --target llama-cli llama-completion llama-server llama-tokenize llama-imatrix \
           llama-quantize llama-perplexity
```

For a CPU-only conversion build, omit `-DGGML_CUDA=ON` and
`-DCMAKE_CUDA_ARCHITECTURES=native`.

## 2. Download and convert the immutable source revision

```bash
hf download inclusionAI/Ling-3.0-tiny \
  --revision a2ee06c0f2de5b171701aee7f73f70a1da75483b \
  --local-dir Ling-3.0-tiny

(cd Ling-3.0-tiny && sha256sum -c ../source-safetensors.sha256)

python llama.cpp/convert_hf_to_gguf.py Ling-3.0-tiny \
  --outtype bf16 \
  --outfile Ling-3.0-tiny-BF16.gguf \
  --verbose
```

The expected BF16 result is 526 tensors, 15,803,475,232 bytes, with SHA-256
`2020d58d44887c4078c310dad0363b2af0898a9ed6cc46675c0d23201833952e`.

## 3. Collect the importance matrix

```bash
python -m pip install duckdb==1.3.2

hf download lemon07r/bartowski-imatrix-v5-semantic \
  bartowski-imatrix-v5-semantic.txt \
  --repo-type dataset \
  --revision a306f203ee4323e0afe846ae02c2daafe17384d9 \
  --local-dir calibration

hf download eaddario/imatrix-calibration \
  combined_all_micro.parquet \
  --repo-type dataset \
  --revision e87ed55dcba9d9c3a3e41539f3e728e981b1daa4 \
  --local-dir calibration/eaddario

python - <<'PY'
import duckdb

source = "calibration/eaddario/combined_all_micro.parquet"
output = "calibration/eaddario/combined_all_micro.txt"
rows = duckdb.connect().execute(
    "SELECT content FROM read_parquet(?)", [source]
).fetchall()
with open(output, "w", encoding="utf-8", newline="\n") as handle:
    for (content,) in rows:
        if content:
            handle.write(content.replace("\x00", ""))
            handle.write("\n")
PY

printf '%s\n' \
  'ff879b5a748f822ef539e43c596a3f44ab922f0295ee209d4220d9f86e86a063  calibration/bartowski-imatrix-v5-semantic.txt' \
  '94389921e1f67b180a99de28c3090b41ce6f1960eb13abad21b7eba7cbe11b26  calibration/eaddario/combined_all_micro.parquet' \
  'fdb2d41abf04a2fb207502741a561a5a9ab385eb0c44a450eae676c410955946  calibration/eaddario/combined_all_micro.txt' \
  | sha256sum -c -

llama.cpp/build/bin/llama-imatrix \
  --model Ling-3.0-tiny-BF16.gguf \
  --file calibration/bartowski-imatrix-v5-semantic.txt \
  --output-file Ling-3.0-tiny-imatrix-primary.gguf \
  --output-format gguf \
  --gpu-layers all --ctx-size 4096 --batch-size 4096 --ubatch-size 512 \
  --threads 16 --threads-batch 16 --device CUDA0 \
  --split-mode none --main-gpu 0 --fit off --flash-attn off \
  --offline --no-ppl \
  --output-frequency 25

llama.cpp/build/bin/llama-imatrix \
  --model Ling-3.0-tiny-BF16.gguf \
  --file calibration/eaddario/combined_all_micro.txt \
  --in-file Ling-3.0-tiny-imatrix-primary.gguf \
  --output-file Ling-3.0-tiny-imatrix.gguf \
  --output-format gguf \
  --gpu-layers all --ctx-size 4096 --batch-size 4096 --ubatch-size 512 \
  --threads 16 --threads-batch 16 --device CUDA0 \
  --split-mode none --main-gpu 0 --fit off --flash-attn off \
  --offline --no-ppl \
  --output-frequency 25
```

The primary text SHA-256 is
`ff879b5a748f822ef539e43c596a3f44ab922f0295ee209d4220d9f86e86a063`.
The supplement parquet and extracted text SHA-256 values are respectively
`94389921e1f67b180a99de28c3090b41ce6f1960eb13abad21b7eba7cbe11b26`
and `fdb2d41abf04a2fb207502741a561a5a9ab385eb0c44a450eae676c410955946`.
They are used only for activation statistics, not training or evaluation.
`--process-output` is intentionally omitted. The pinned tool documentation
states that it is typically better not to use the importance matrix for
`output.weight`, which is why collection for that tensor defaults to false.

The commands above are tensor-equivalent reproduction commands. The released
matrix and importance-aware model files also embed path strings. Byte-for-byte
SHA-256 reproduction requires the primary dataset at
`/workspace/ling3/calibration/bartowski-imatrix-v5-semantic.txt`, the extracted
supplement at `/workspace/ling3/calibration/eaddario/combined_all_micro.txt`,
and the final matrix at `/tmp/Ling-3.0-tiny-imatrix.gguf`. The source directory
basename must be `Ling-3.0-tiny`. Different path spellings change metadata
bytes without changing the collected statistics or quantized tensor values.

The released final matrix is 44,016,768 bytes with SHA-256
`e8b15d131f9ce294f922c5c387f7a69829c12100d6a35bb1635a2b859083c3f0`.
It contains 332 entries from 162 complete chunks, and all 8,832 routed-expert
count slots are nonzero.

## 4. Quantize directly from BF16

The matrix is deliberately used for K-quants as well as IQ quants. No output
is requantized from Q8 or another reduced-precision artifact.

```bash
llama.cpp/build/bin/llama-quantize \
  Ling-3.0-tiny-BF16.gguf Ling-3.0-tiny-Q8_0.gguf Q8_0 32

for quant in Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
  llama.cpp/build/bin/llama-quantize \
    --imatrix Ling-3.0-tiny-imatrix.gguf \
    Ling-3.0-tiny-BF16.gguf "Ling-3.0-tiny-${quant}.gguf" "$quant" 32
done
```

Do not add `--allow-requantize` or `--pure` when reproducing these files.
IQ4_NL and MXFP4_MOE candidates were measured but intentionally rejected; they
are not part of the published artifact set.

## 5. Held-out PPL and KLD comparison

Validation uses WikiText-2 from `ggml-org/ci` at revision
`927b3642933080f1b0e811e2f916e14c292992f9`. The extracted test file SHA-256
is `173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08`.

```bash
hf download ggml-org/ci wikitext-2-raw-v1.zip \
  --repo-type dataset \
  --revision 927b3642933080f1b0e811e2f916e14c292992f9 \
  --local-dir validation
unzip -q validation/wikitext-2-raw-v1.zip -d validation

printf '%s\n' \
  'ef7edb566e3e2b2d31b29c1fdb0c89a4cc683597484c3dc2517919c615435a11  validation/wikitext-2-raw-v1.zip' \
  '173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08  validation/wikitext-2-raw/wiki.test.raw' \
  | sha256sum -c -
```

The archive SHA-256 is
`ef7edb566e3e2b2d31b29c1fdb0c89a4cc683597484c3dc2517919c615435a11`.

First create the BF16 log-probability reference. In this pinned build,
`--kl-divergence-base` is an alias for `--save-all-logits`.

```bash
llama.cpp/build/bin/llama-perplexity \
  --model Ling-3.0-tiny-BF16.gguf \
  --file validation/wikitext-2-raw/wiki.test.raw \
  --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
  --chunks 32 --threads 16 --threads-batch 16 --device CUDA0 \
  --split-mode none --main-gpu 0 --fit off --flash-attn off \
  --kv-offload --op-offload --no-repack \
  --cache-type-k f16 --cache-type-v f16 \
  --no-warmup --offline \
  --kl-divergence-base validation/bf16-c512-chunks32.kld
```

Then compare each quant using the same context and batch sizes. The comparator
reads the exact tokens and chunk count from the reference file.

```bash
llama.cpp/build/bin/llama-perplexity \
  --model Ling-3.0-tiny-BF16.gguf \
  --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
  --threads 16 --threads-batch 16 --device CUDA0 \
  --split-mode none --main-gpu 0 --fit off --flash-attn off \
  --kv-offload --op-offload --no-repack \
  --cache-type-k f16 --cache-type-v f16 \
  --no-warmup --offline \
  --kl-divergence-base validation/bf16-c512-chunks32.kld \
  --kl-divergence \
  > validation/kld-BF16.log 2>&1

for quant in Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
  llama.cpp/build/bin/llama-perplexity \
    --model "Ling-3.0-tiny-${quant}.gguf" \
    --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
    --threads 16 --threads-batch 16 --device CUDA0 \
    --split-mode none --main-gpu 0 --fit off --flash-attn off \
    --kv-offload --op-offload --no-repack \
    --cache-type-k f16 --cache-type-v f16 \
    --no-warmup --offline \
    --kl-divergence-base validation/bf16-c512-chunks32.kld \
    --kl-divergence \
    > "validation/kld-${quant}.log" 2>&1
done

python validation/parse_kld.py --output validation/kld-results.json \
  BF16=validation/kld-BF16.log \
  Q8_0=validation/kld-Q8_0.log \
  Q6_K=validation/kld-Q6_K.log \
  Q5_K_M=validation/kld-Q5_K_M.log \
  Q4_K_M=validation/kld-Q4_K_M.log \
  Q4_K_S=validation/kld-Q4_K_S.log \
  IQ4_XS=validation/kld-IQ4_XS.log \
  Q3_K_M=validation/kld-Q3_K_M.log \
  IQ3_M=validation/kld-IQ3_M.log \
  IQ2_M=validation/kld-IQ2_M.log
```

Do not combine this KLD workflow with `--ppl-stride`; it uses a different
evaluation path. Comparison logs must be newly truncated, never appended. The
parser requires a complete terminal statistics block because error paths in
this pinned `llama-perplexity` can still return process status zero.

The stored BF16 reference is 2,565,373,716 bytes with SHA-256
`afa4dc9bd2d995dd85d8614494dc9cccec8e0e4bd6b2c24efc81d3bba96638f7`.

## 6. Structure, tokenizer, and deterministic generation

Generate a canonical shape map from BF16, then compare every quant against it.
The validator also checks the architecture metadata, Q-LoRA tensors, exact
embedded chat template, NEXTN/MTP absence, matrix provenance, and the exact
filename-bound tensor-type inventory of every published model. Its separate
MXFP4 whitelist is retained for auditing rejected candidates.

```bash
python validation/check_gguf_structure.py \
  --gguf-dump .venv/bin/gguf-dump \
  --model Ling-3.0-tiny-BF16.gguf \
  --source-template Ling-3.0-tiny/chat_template.jinja \
  --write-reference-shapes validation/bf16-tensor-shapes.json \
  --output validation/structure-BF16.json

for quant in Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
  python validation/check_gguf_structure.py \
    --gguf-dump .venv/bin/gguf-dump \
    --model "Ling-3.0-tiny-${quant}.gguf" \
    --source-template Ling-3.0-tiny/chat_template.jinja \
    --reference-shapes validation/bf16-tensor-shapes.json \
    --output "validation/structure-${quant}.json"
done
```

Generate the independent Transformers reference only after verifying all 32
source shards and the pinned control files:

```bash
python3 -m venv .hfref-venv
.hfref-venv/bin/python -m pip install \
  --extra-index-url https://download.pytorch.org/whl/cu128 \
  accelerate==1.10.1 einops==0.8.1 fla-core==0.5.1 \
  huggingface_hub==0.36.2 numpy==2.1.2 safetensors==0.8.0 \
  tokenizers==0.22.2 'torch==2.8.0+cu128' tqdm==4.70.0 \
  transformers==4.57.6

.hfref-venv/bin/python validation/hf_reference.py \
  Ling-3.0-tiny validation/hf-reference.json \
  --source-weight-manifest source-safetensors.sha256

python validation/compare_tokenizer.py \
  --llama-tokenize llama.cpp/build/bin/llama-tokenize \
  --model Ling-3.0-tiny-BF16.gguf \
  --hf-reference validation/hf-reference.json \
  --output validation/tokenizer-comparison.json
```

The controlled raw-generation smoke uses the exact same full prompt in both
runtimes. `compare_generation.py` now fails if the supplied prompt differs
from the reference, even when the generated suffix happens to match.

```bash
prompt='The capital of France is Paris. The capital of Germany is Berlin. The capital of Japan is'

for quant in BF16 Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
  llama.cpp/build/bin/llama-completion \
    --model "Ling-3.0-tiny-${quant}.gguf" \
    --prompt "$prompt" --predict 12 \
    --ctx-size 512 --batch-size 512 --ubatch-size 128 \
    --threads 16 --threads-batch 16 --device CUDA0 --gpu-layers all \
    --split-mode none --main-gpu 0 --fit off --flash-attn off \
    --kv-offload --op-offload --no-repack \
    --cache-type-k f16 --cache-type-v f16 \
    --seed 1 --samplers temperature --temperature 0 \
    -no-cnv --no-context-shift --no-warmup --no-display-prompt \
    --color off --offline --log-colors off --no-log-timestamps \
    --check-tensors \
    > "validation/gguf-greedy-${quant}.txt" \
    2> "validation/load-${quant}.log"

  python validation/compare_generation.py \
    --artifact "$quant" \
    --hf-reference validation/hf-reference.json \
    --gguf-output "validation/gguf-greedy-${quant}.txt" \
    --case greedy_capitals_12 --prompt-text "$prompt" \
    --mismatch-policy fail \
    --output "validation/generation-${quant}.json"
done
```

## 7. Multiple-choice collapse and 32K execution screens

```bash
hf download ikawrakow/validation-datasets-for-llama.cpp \
  mmlu-validation.bin --repo-type dataset \
  --revision 37884b81b4957f1950a53b6ff48d77c8dd5e430c \
  --local-dir validation

printf '%s\n' \
  '470af3a74eccacfaf6f43b08aabf510f61e6c92fe20d17241ded934151e225fa  validation/mmlu-validation.bin' \
  | sha256sum -c -

for quant in BF16 Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
  llama.cpp/build/bin/llama-perplexity \
    --model "Ling-3.0-tiny-${quant}.gguf" \
    --file validation/mmlu-validation.bin \
    --multiple-choice --multiple-choice-tasks 500 --seed 1 \
    --gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
    --threads 16 --threads-batch 16 --device CUDA0 \
    --split-mode none --main-gpu 0 --fit off --flash-attn off \
    --kv-offload --op-offload --no-repack \
    --cache-type-k f16 --cache-type-v f16 \
    --no-warmup --offline --no-log-timestamps \
    > "validation/mmlu-${quant}.log" 2>&1
done
```

The pinned tool prints “TruthfulQA” internally, but the supplied binary is the
pinned MMLU validation file. Results are summarized in
`validation/multiple-choice-results.json`.

For BF16, Q4_K_M, and IQ2_M, run the 32K execution/prefill screen as follows:

```bash
for quant in BF16 Q4_K_M IQ2_M; do
  llama.cpp/build/bin/llama-perplexity \
    --model "Ling-3.0-tiny-${quant}.gguf" \
    --file validation/wikitext-2-raw/wiki.test.raw \
    --chunks 1 --ctx-size 32768 --batch-size 4096 --ubatch-size 512 \
    --gpu-layers all --threads 16 --threads-batch 16 --device CUDA0 \
    --split-mode none --main-gpu 0 --fit off --flash-attn off \
    --kv-offload --op-offload --no-repack \
    --cache-type-k f16 --cache-type-v f16 \
    --no-warmup --offline --no-log-timestamps \
    > "validation/context-32768-${quant}.log" 2>&1
done
```

This is not exhaustive validation of the native 131K limit.

## 8. Server/Jinja and final integrity checks

The exact four OpenAI-compatible server requests and their expected semantic
checks are stored in `validation/server-requests.json` and
`validation/server-results.json`. Start Q4_K_M with `llama-server --jinja`,
submit each request to `/v1/chat/completions`, and verify thinking on/off,
Chinese generation, normal EOS stops, and the two-argument required tool call.

```bash
llama.cpp/build/bin/llama-server \
  --model Ling-3.0-tiny-Q4_K_M.gguf --alias ling-3.0-tiny \
  --host 127.0.0.1 --port 8080 --jinja --ctx-size 8192 \
  --gpu-layers 999 --split-mode none --main-gpu 0 --device CUDA0 \
  --flash-attn off --no-warmup --offline

# In another terminal after the server reports that it is listening:
python validation/run_server_requests.py \
  --requests validation/server-requests.json \
  --response-dir validation
python validation/validate_server_responses.py \
  --requests validation/server-requests.json \
  --response-dir validation \
  --output validation/server-results.json
```

Finally, verify every downloaded release artifact from the repository root:

```bash
sha256sum -c SHA256SUMS
```

`conversion_manifest.json` records the exact source/control hashes, build and
binary hashes, package versions, hardware, calibration order, artifact tensor
inventories, Hub introduction commits, and validation-report bindings.