Instructions to use Mike0021/Ling-3.0-tiny-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Mike0021/Ling-3.0-tiny-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Mike0021/Ling-3.0-tiny-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mike0021/Ling-3.0-tiny-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mike0021/Ling-3.0-tiny-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Ollama
How to use Mike0021/Ling-3.0-tiny-GGUF with Ollama:
ollama run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Unsloth Studio
How to use Mike0021/Ling-3.0-tiny-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Mike0021/Ling-3.0-tiny-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Mike0021/Ling-3.0-tiny-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Mike0021/Ling-3.0-tiny-GGUF to start chatting
- Pi
How to use Mike0021/Ling-3.0-tiny-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use Mike0021/Ling-3.0-tiny-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Mike0021/Ling-3.0-tiny-GGUF with Docker Model Runner:
docker model run hf.co/Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
- Lemonade
How to use Mike0021/Ling-3.0-tiny-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ling-3.0-tiny-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Mike0021/Ling-3.0-tiny-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Mike0021/Ling-3.0-tiny-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
File size: 16,305 Bytes
31596cd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 | # Reproducing this GGUF release
The release is pinned to immutable source, converter, calibration, and
validation revisions. BailingMoE3 support is not yet merged in upstream
`llama.cpp`; use the exact revision below.
## 1. Build the converter and runtime
```bash
git clone https://github.com/aetherbird/llama.cpp.git
git -C llama.cpp checkout d8d862521e9ad842f2b47f3b392b039317782aa0
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -r llama.cpp/requirements.txt
python -m pip install --extra-index-url https://download.pytorch.org/whl/cu128 \
gguf==0.19.0 huggingface_hub==0.36.2 numpy==1.26.4 \
protobuf==4.25.9 safetensors==0.8.0 sentencepiece==0.2.2 \
tokenizers==0.22.2 'torch==2.8.0+cu128' tqdm==4.70.0 \
transformers==4.57.6
cmake -S llama.cpp -B llama.cpp/build \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_CUDA=ON \
-DGGML_NATIVE=OFF \
-DCMAKE_CUDA_ARCHITECTURES=native
cmake --build llama.cpp/build --config Release --parallel \
--target llama-cli llama-completion llama-server llama-tokenize llama-imatrix \
llama-quantize llama-perplexity
```
For a CPU-only conversion build, omit `-DGGML_CUDA=ON` and
`-DCMAKE_CUDA_ARCHITECTURES=native`.
## 2. Download and convert the immutable source revision
```bash
hf download inclusionAI/Ling-3.0-tiny \
--revision a2ee06c0f2de5b171701aee7f73f70a1da75483b \
--local-dir Ling-3.0-tiny
(cd Ling-3.0-tiny && sha256sum -c ../source-safetensors.sha256)
python llama.cpp/convert_hf_to_gguf.py Ling-3.0-tiny \
--outtype bf16 \
--outfile Ling-3.0-tiny-BF16.gguf \
--verbose
```
The expected BF16 result is 526 tensors, 15,803,475,232 bytes, with SHA-256
`2020d58d44887c4078c310dad0363b2af0898a9ed6cc46675c0d23201833952e`.
## 3. Collect the importance matrix
```bash
python -m pip install duckdb==1.3.2
hf download lemon07r/bartowski-imatrix-v5-semantic \
bartowski-imatrix-v5-semantic.txt \
--repo-type dataset \
--revision a306f203ee4323e0afe846ae02c2daafe17384d9 \
--local-dir calibration
hf download eaddario/imatrix-calibration \
combined_all_micro.parquet \
--repo-type dataset \
--revision e87ed55dcba9d9c3a3e41539f3e728e981b1daa4 \
--local-dir calibration/eaddario
python - <<'PY'
import duckdb
source = "calibration/eaddario/combined_all_micro.parquet"
output = "calibration/eaddario/combined_all_micro.txt"
rows = duckdb.connect().execute(
"SELECT content FROM read_parquet(?)", [source]
).fetchall()
with open(output, "w", encoding="utf-8", newline="\n") as handle:
for (content,) in rows:
if content:
handle.write(content.replace("\x00", ""))
handle.write("\n")
PY
printf '%s\n' \
'ff879b5a748f822ef539e43c596a3f44ab922f0295ee209d4220d9f86e86a063 calibration/bartowski-imatrix-v5-semantic.txt' \
'94389921e1f67b180a99de28c3090b41ce6f1960eb13abad21b7eba7cbe11b26 calibration/eaddario/combined_all_micro.parquet' \
'fdb2d41abf04a2fb207502741a561a5a9ab385eb0c44a450eae676c410955946 calibration/eaddario/combined_all_micro.txt' \
| sha256sum -c -
llama.cpp/build/bin/llama-imatrix \
--model Ling-3.0-tiny-BF16.gguf \
--file calibration/bartowski-imatrix-v5-semantic.txt \
--output-file Ling-3.0-tiny-imatrix-primary.gguf \
--output-format gguf \
--gpu-layers all --ctx-size 4096 --batch-size 4096 --ubatch-size 512 \
--threads 16 --threads-batch 16 --device CUDA0 \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--offline --no-ppl \
--output-frequency 25
llama.cpp/build/bin/llama-imatrix \
--model Ling-3.0-tiny-BF16.gguf \
--file calibration/eaddario/combined_all_micro.txt \
--in-file Ling-3.0-tiny-imatrix-primary.gguf \
--output-file Ling-3.0-tiny-imatrix.gguf \
--output-format gguf \
--gpu-layers all --ctx-size 4096 --batch-size 4096 --ubatch-size 512 \
--threads 16 --threads-batch 16 --device CUDA0 \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--offline --no-ppl \
--output-frequency 25
```
The primary text SHA-256 is
`ff879b5a748f822ef539e43c596a3f44ab922f0295ee209d4220d9f86e86a063`.
The supplement parquet and extracted text SHA-256 values are respectively
`94389921e1f67b180a99de28c3090b41ce6f1960eb13abad21b7eba7cbe11b26`
and `fdb2d41abf04a2fb207502741a561a5a9ab385eb0c44a450eae676c410955946`.
They are used only for activation statistics, not training or evaluation.
`--process-output` is intentionally omitted. The pinned tool documentation
states that it is typically better not to use the importance matrix for
`output.weight`, which is why collection for that tensor defaults to false.
The commands above are tensor-equivalent reproduction commands. The released
matrix and importance-aware model files also embed path strings. Byte-for-byte
SHA-256 reproduction requires the primary dataset at
`/workspace/ling3/calibration/bartowski-imatrix-v5-semantic.txt`, the extracted
supplement at `/workspace/ling3/calibration/eaddario/combined_all_micro.txt`,
and the final matrix at `/tmp/Ling-3.0-tiny-imatrix.gguf`. The source directory
basename must be `Ling-3.0-tiny`. Different path spellings change metadata
bytes without changing the collected statistics or quantized tensor values.
The released final matrix is 44,016,768 bytes with SHA-256
`e8b15d131f9ce294f922c5c387f7a69829c12100d6a35bb1635a2b859083c3f0`.
It contains 332 entries from 162 complete chunks, and all 8,832 routed-expert
count slots are nonzero.
## 4. Quantize directly from BF16
The matrix is deliberately used for K-quants as well as IQ quants. No output
is requantized from Q8 or another reduced-precision artifact.
```bash
llama.cpp/build/bin/llama-quantize \
Ling-3.0-tiny-BF16.gguf Ling-3.0-tiny-Q8_0.gguf Q8_0 32
for quant in Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
llama.cpp/build/bin/llama-quantize \
--imatrix Ling-3.0-tiny-imatrix.gguf \
Ling-3.0-tiny-BF16.gguf "Ling-3.0-tiny-${quant}.gguf" "$quant" 32
done
```
Do not add `--allow-requantize` or `--pure` when reproducing these files.
IQ4_NL and MXFP4_MOE candidates were measured but intentionally rejected; they
are not part of the published artifact set.
## 5. Held-out PPL and KLD comparison
Validation uses WikiText-2 from `ggml-org/ci` at revision
`927b3642933080f1b0e811e2f916e14c292992f9`. The extracted test file SHA-256
is `173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08`.
```bash
hf download ggml-org/ci wikitext-2-raw-v1.zip \
--repo-type dataset \
--revision 927b3642933080f1b0e811e2f916e14c292992f9 \
--local-dir validation
unzip -q validation/wikitext-2-raw-v1.zip -d validation
printf '%s\n' \
'ef7edb566e3e2b2d31b29c1fdb0c89a4cc683597484c3dc2517919c615435a11 validation/wikitext-2-raw-v1.zip' \
'173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08 validation/wikitext-2-raw/wiki.test.raw' \
| sha256sum -c -
```
The archive SHA-256 is
`ef7edb566e3e2b2d31b29c1fdb0c89a4cc683597484c3dc2517919c615435a11`.
First create the BF16 log-probability reference. In this pinned build,
`--kl-divergence-base` is an alias for `--save-all-logits`.
```bash
llama.cpp/build/bin/llama-perplexity \
--model Ling-3.0-tiny-BF16.gguf \
--file validation/wikitext-2-raw/wiki.test.raw \
--gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
--chunks 32 --threads 16 --threads-batch 16 --device CUDA0 \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--kv-offload --op-offload --no-repack \
--cache-type-k f16 --cache-type-v f16 \
--no-warmup --offline \
--kl-divergence-base validation/bf16-c512-chunks32.kld
```
Then compare each quant using the same context and batch sizes. The comparator
reads the exact tokens and chunk count from the reference file.
```bash
llama.cpp/build/bin/llama-perplexity \
--model Ling-3.0-tiny-BF16.gguf \
--gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
--threads 16 --threads-batch 16 --device CUDA0 \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--kv-offload --op-offload --no-repack \
--cache-type-k f16 --cache-type-v f16 \
--no-warmup --offline \
--kl-divergence-base validation/bf16-c512-chunks32.kld \
--kl-divergence \
> validation/kld-BF16.log 2>&1
for quant in Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
llama.cpp/build/bin/llama-perplexity \
--model "Ling-3.0-tiny-${quant}.gguf" \
--gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
--threads 16 --threads-batch 16 --device CUDA0 \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--kv-offload --op-offload --no-repack \
--cache-type-k f16 --cache-type-v f16 \
--no-warmup --offline \
--kl-divergence-base validation/bf16-c512-chunks32.kld \
--kl-divergence \
> "validation/kld-${quant}.log" 2>&1
done
python validation/parse_kld.py --output validation/kld-results.json \
BF16=validation/kld-BF16.log \
Q8_0=validation/kld-Q8_0.log \
Q6_K=validation/kld-Q6_K.log \
Q5_K_M=validation/kld-Q5_K_M.log \
Q4_K_M=validation/kld-Q4_K_M.log \
Q4_K_S=validation/kld-Q4_K_S.log \
IQ4_XS=validation/kld-IQ4_XS.log \
Q3_K_M=validation/kld-Q3_K_M.log \
IQ3_M=validation/kld-IQ3_M.log \
IQ2_M=validation/kld-IQ2_M.log
```
Do not combine this KLD workflow with `--ppl-stride`; it uses a different
evaluation path. Comparison logs must be newly truncated, never appended. The
parser requires a complete terminal statistics block because error paths in
this pinned `llama-perplexity` can still return process status zero.
The stored BF16 reference is 2,565,373,716 bytes with SHA-256
`afa4dc9bd2d995dd85d8614494dc9cccec8e0e4bd6b2c24efc81d3bba96638f7`.
## 6. Structure, tokenizer, and deterministic generation
Generate a canonical shape map from BF16, then compare every quant against it.
The validator also checks the architecture metadata, Q-LoRA tensors, exact
embedded chat template, NEXTN/MTP absence, matrix provenance, and the exact
filename-bound tensor-type inventory of every published model. Its separate
MXFP4 whitelist is retained for auditing rejected candidates.
```bash
python validation/check_gguf_structure.py \
--gguf-dump .venv/bin/gguf-dump \
--model Ling-3.0-tiny-BF16.gguf \
--source-template Ling-3.0-tiny/chat_template.jinja \
--write-reference-shapes validation/bf16-tensor-shapes.json \
--output validation/structure-BF16.json
for quant in Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
python validation/check_gguf_structure.py \
--gguf-dump .venv/bin/gguf-dump \
--model "Ling-3.0-tiny-${quant}.gguf" \
--source-template Ling-3.0-tiny/chat_template.jinja \
--reference-shapes validation/bf16-tensor-shapes.json \
--output "validation/structure-${quant}.json"
done
```
Generate the independent Transformers reference only after verifying all 32
source shards and the pinned control files:
```bash
python3 -m venv .hfref-venv
.hfref-venv/bin/python -m pip install \
--extra-index-url https://download.pytorch.org/whl/cu128 \
accelerate==1.10.1 einops==0.8.1 fla-core==0.5.1 \
huggingface_hub==0.36.2 numpy==2.1.2 safetensors==0.8.0 \
tokenizers==0.22.2 'torch==2.8.0+cu128' tqdm==4.70.0 \
transformers==4.57.6
.hfref-venv/bin/python validation/hf_reference.py \
Ling-3.0-tiny validation/hf-reference.json \
--source-weight-manifest source-safetensors.sha256
python validation/compare_tokenizer.py \
--llama-tokenize llama.cpp/build/bin/llama-tokenize \
--model Ling-3.0-tiny-BF16.gguf \
--hf-reference validation/hf-reference.json \
--output validation/tokenizer-comparison.json
```
The controlled raw-generation smoke uses the exact same full prompt in both
runtimes. `compare_generation.py` now fails if the supplied prompt differs
from the reference, even when the generated suffix happens to match.
```bash
prompt='The capital of France is Paris. The capital of Germany is Berlin. The capital of Japan is'
for quant in BF16 Q8_0 Q6_K Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
llama.cpp/build/bin/llama-completion \
--model "Ling-3.0-tiny-${quant}.gguf" \
--prompt "$prompt" --predict 12 \
--ctx-size 512 --batch-size 512 --ubatch-size 128 \
--threads 16 --threads-batch 16 --device CUDA0 --gpu-layers all \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--kv-offload --op-offload --no-repack \
--cache-type-k f16 --cache-type-v f16 \
--seed 1 --samplers temperature --temperature 0 \
-no-cnv --no-context-shift --no-warmup --no-display-prompt \
--color off --offline --log-colors off --no-log-timestamps \
--check-tensors \
> "validation/gguf-greedy-${quant}.txt" \
2> "validation/load-${quant}.log"
python validation/compare_generation.py \
--artifact "$quant" \
--hf-reference validation/hf-reference.json \
--gguf-output "validation/gguf-greedy-${quant}.txt" \
--case greedy_capitals_12 --prompt-text "$prompt" \
--mismatch-policy fail \
--output "validation/generation-${quant}.json"
done
```
## 7. Multiple-choice collapse and 32K execution screens
```bash
hf download ikawrakow/validation-datasets-for-llama.cpp \
mmlu-validation.bin --repo-type dataset \
--revision 37884b81b4957f1950a53b6ff48d77c8dd5e430c \
--local-dir validation
printf '%s\n' \
'470af3a74eccacfaf6f43b08aabf510f61e6c92fe20d17241ded934151e225fa validation/mmlu-validation.bin' \
| sha256sum -c -
for quant in BF16 Q5_K_M Q4_K_M Q4_K_S IQ4_XS Q3_K_M IQ3_M IQ2_M; do
llama.cpp/build/bin/llama-perplexity \
--model "Ling-3.0-tiny-${quant}.gguf" \
--file validation/mmlu-validation.bin \
--multiple-choice --multiple-choice-tasks 500 --seed 1 \
--gpu-layers all --ctx-size 512 --batch-size 512 --ubatch-size 512 \
--threads 16 --threads-batch 16 --device CUDA0 \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--kv-offload --op-offload --no-repack \
--cache-type-k f16 --cache-type-v f16 \
--no-warmup --offline --no-log-timestamps \
> "validation/mmlu-${quant}.log" 2>&1
done
```
The pinned tool prints “TruthfulQA” internally, but the supplied binary is the
pinned MMLU validation file. Results are summarized in
`validation/multiple-choice-results.json`.
For BF16, Q4_K_M, and IQ2_M, run the 32K execution/prefill screen as follows:
```bash
for quant in BF16 Q4_K_M IQ2_M; do
llama.cpp/build/bin/llama-perplexity \
--model "Ling-3.0-tiny-${quant}.gguf" \
--file validation/wikitext-2-raw/wiki.test.raw \
--chunks 1 --ctx-size 32768 --batch-size 4096 --ubatch-size 512 \
--gpu-layers all --threads 16 --threads-batch 16 --device CUDA0 \
--split-mode none --main-gpu 0 --fit off --flash-attn off \
--kv-offload --op-offload --no-repack \
--cache-type-k f16 --cache-type-v f16 \
--no-warmup --offline --no-log-timestamps \
> "validation/context-32768-${quant}.log" 2>&1
done
```
This is not exhaustive validation of the native 131K limit.
## 8. Server/Jinja and final integrity checks
The exact four OpenAI-compatible server requests and their expected semantic
checks are stored in `validation/server-requests.json` and
`validation/server-results.json`. Start Q4_K_M with `llama-server --jinja`,
submit each request to `/v1/chat/completions`, and verify thinking on/off,
Chinese generation, normal EOS stops, and the two-argument required tool call.
```bash
llama.cpp/build/bin/llama-server \
--model Ling-3.0-tiny-Q4_K_M.gguf --alias ling-3.0-tiny \
--host 127.0.0.1 --port 8080 --jinja --ctx-size 8192 \
--gpu-layers 999 --split-mode none --main-gpu 0 --device CUDA0 \
--flash-attn off --no-warmup --offline
# In another terminal after the server reports that it is listening:
python validation/run_server_requests.py \
--requests validation/server-requests.json \
--response-dir validation
python validation/validate_server_responses.py \
--requests validation/server-requests.json \
--response-dir validation \
--output validation/server-results.json
```
Finally, verify every downloaded release artifact from the repository root:
```bash
sha256sum -c SHA256SUMS
```
`conversion_manifest.json` records the exact source/control hashes, build and
binary hashes, package versions, hardware, calibration order, artifact tensor
inventories, Hub introduction commits, and validation-report bindings.
|