Build and verification record
Objective
Create a reproducible DS4 package that maximizes target-model quality within a 128 GB unified-memory envelope. Preserve exact native MXFP4 routed experts on the ten sensitivity-selected layers and retain the proven low-bit/Q8 policy everywhere else.
Inputs and toolchain
- All weights were regenerated from
apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8. - The sparse reference GGUF supplied metadata, tokenizer, tensor order and shapes only.
- All 48 source-shard hashes, source metadata and the abliteration manifest were verified before conversion.
- The routed-expert imatrix supplied exact coverage for the 129 tensors whose output types consume calibration data.
- DS4 source revision:
d516d4eeb82c454aeb2831af1b1961801d6b571b. - Upstream
ds4f-mxfp4base:4893e0c40fba03dbc85555faeb035799aa04e0b6. - Quantizer SHA-256:
f0a381f4ada808ea2afa740d964354fa327fc1235ba7cebf50874eb89fb97ac5. - Runtime SHA-256:
2aaf20469b9918d6d6ab8787a02811c11228547cd979787879a03dba8a9e7824.
Quantization policy
- Native MXFP4 gate/up/down routed experts on layers
10, 14, 30, 34, 37, 38, 39, 40, 41, 42. - IQ2_XXS gate/up and Q2_K down routed experts on the other 33 MoE layers.
- Q8 attention, shared-expert and output tensors.
- Protected F16 indexer and auxiliary tensors.
- DSpark support with IQ2_XXS gate/up and native MXFP4 down projections for target layers 40, 41 and 42.
Conversion gates
- Strict imatrix mode passed for main and DSpark conversions.
- Main GGUF contains exactly 1,328 tensors and the intended 30 MXFP4 tensors.
- DSpark support contains exactly 81 tensors and three MXFP4 down aggregates.
- Every native tensor comparison required an explicit
byte_compare: OK. - All 30 main and three DSpark MXFP4 tensors independently reproduced the source FP4 codes and scale bytes.
ds4 --cpu --inspect --dspark-strictreported zero missing tensors, invalid bindings or metadata errors.- The final GGUF hashes are recorded in
SHA256SUMS.
Result
| Component | Bytes | GiB |
|---|---|---|
| Main GGUF | 102,826,238,912 |
95.7644 |
| DSpark support | 7,297,737,120 |
6.7965 |
| Combined | 110,123,976,032 |
102.5609 |
Native MXFP4 improves fidelity over a second Q4_K requantization and uses 4.25 bits per weight rather than 4.50. Native MXFP4 Metal and the mixed DSpark path are newer than mature Q4_K kernels, so throughput should be measured on the target system.
One-million-token context
Recommended target-only mode:
ds4 --metal \
-m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
--ctx 1048576 --prefill-chunk 2048
Estimated residency is at most 110.30 GiB. DSpark can be tested with a 1,024
token prefill chunk; estimated residency is 114.02 GiB. A 4,096-token chunk
with DSpark is estimated at 123.24 GiB and is not a reliably resident mode on
a machine whose Metal recommended working set is approximately 121.60 GiB.