# Build and verification record ## Objective Create a reproducible DS4 package that maximizes target-model quality within a 128 GB unified-memory envelope. Preserve exact native MXFP4 routed experts on the ten sensitivity-selected layers and retain the proven low-bit/Q8 policy everywhere else. ## Inputs and toolchain - All weights were regenerated from `apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8`. - The sparse reference GGUF supplied metadata, tokenizer, tensor order and shapes only. - All 48 source-shard hashes, source metadata and the abliteration manifest were verified before conversion. - The routed-expert imatrix supplied exact coverage for the 129 tensors whose output types consume calibration data. - DS4 source revision: `d516d4eeb82c454aeb2831af1b1961801d6b571b`. - Upstream `ds4f-mxfp4` base: `4893e0c40fba03dbc85555faeb035799aa04e0b6`. - Quantizer SHA-256: `f0a381f4ada808ea2afa740d964354fa327fc1235ba7cebf50874eb89fb97ac5`. - Runtime SHA-256: `2aaf20469b9918d6d6ab8787a02811c11228547cd979787879a03dba8a9e7824`. ## Quantization policy - Native MXFP4 gate/up/down routed experts on layers `10, 14, 30, 34, 37, 38, 39, 40, 41, 42`. - IQ2_XXS gate/up and Q2_K down routed experts on the other 33 MoE layers. - Q8 attention, shared-expert and output tensors. - Protected F16 indexer and auxiliary tensors. - DSpark support with IQ2_XXS gate/up and native MXFP4 down projections for target layers 40, 41 and 42. ## Conversion gates - Strict imatrix mode passed for main and DSpark conversions. - Main GGUF contains exactly 1,328 tensors and the intended 30 MXFP4 tensors. - DSpark support contains exactly 81 tensors and three MXFP4 down aggregates. - Every native tensor comparison required an explicit `byte_compare: OK`. - All 30 main and three DSpark MXFP4 tensors independently reproduced the source FP4 codes and scale bytes. - `ds4 --cpu --inspect --dspark-strict` reported zero missing tensors, invalid bindings or metadata errors. - The final GGUF hashes are recorded in `SHA256SUMS`. ## Result | Component | Bytes | GiB | | --- | ---: | ---: | | Main GGUF | `102,826,238,912` | `95.7644` | | DSpark support | `7,297,737,120` | `6.7965` | | Combined | `110,123,976,032` | `102.5609` | Native MXFP4 improves fidelity over a second Q4_K requantization and uses 4.25 bits per weight rather than 4.50. Native MXFP4 Metal and the mixed DSpark path are newer than mature Q4_K kernels, so throughput should be measured on the target system. ## One-million-token context Recommended target-only mode: ```sh ds4 --metal \ -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \ --ctx 1048576 --prefill-chunk 2048 ``` Estimated residency is at most `110.30 GiB`. DSpark can be tested with a 1,024 token prefill chunk; estimated residency is `114.02 GiB`. A 4,096-token chunk with DSpark is estimated at `123.24 GiB` and is not a reliably resident mode on a machine whose Metal recommended working set is approximately `121.60 GiB`.