--- license: mit library_name: gguf pipeline_tag: text-generation base_model: deepseek-ai/DeepSeek-V4-Flash-0731 base_model_relation: quantized quantized_by: apetersson tags: - gguf - quantized - deepseek - deepseek-v4 - deepseek-v4-flash - moe - mixture-of-experts - mxfp4 - iq2_xxs - q2_k - ds4 - dspark - apple-silicon - metal --- # DeepSeek V4 Flash 0731 — DS4 Quality128 > **Clean official weights, exact MXFP4 experts, maximum resident quality.** > **Required DS4 version:** this model is not compatible with an older DS4 > build. Its native MXFP4 tensors require > [a recent ds4 version from the main branch](https://github.com/antirez/ds4). > That version runs both the target model and the supplied DSpark support model. > **Validation status:** conversion and CPU structural validation are complete. > Metal generation, DSpark acceptance, throughput, peak-memory and actual > one-million-token-context tests for this exact artifact are still pending. This is a quality-first DS4 package of the official [`deepseek-ai/DeepSeek-V4-Flash-0731`](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731) checkpoint. It is designed to keep the target model—and optionally its real three-stage DSpark speculative drafter—resident on a 128 GB Apple Silicon host. This model was quantized directly from the original official checkpoint. No behavioral weight edit was applied before quantization. All model weights were regenerated from the official local FP8 checkpoint; no tensor values were copied from another GGUF or quantized model. ## Artifacts | File | Purpose | Bytes | GiB | SHA-256 | | --- | --- | ---: | ---: | --- | | `DeepSeek-V4-Flash-0731-DS4-Quality128.gguf` | Authoritative 43-layer target model | `102,826,238,912` | `95.7644` | `efcbf786154c2aec61511785d04a4c8d98ee42500b60ff8a810717fc0a69d1d3` | | `DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf` | Three-stage speculative drafter; not standalone | `7,297,737,120` | `6.7965` | `393b807aa0a409f1c006f0e700bbb91a80d3de28daa56c222edf944250075c66` | | **Combined GGUFs** | — | **`110,123,976,032`** | **`102.5609`** | — | `BUILD_MANIFEST.json`, `PROVENANCE.md` and `SHA256SUMS` provide the full machine-readable build record and integrity inventory. ## Quantization profile ### Main target model | Tensor class | Quantization/storage | Rationale | | --- | --- | --- | | Routed gate/up/down on layers 10, 14, 30, 34, 37, 38, 39, 40, 41 and 42 | Exact native `MXFP4` | Preserves source packed I8 codes and F8_E8M0 scales byte-for-byte on sensitivity-selected layers. | | Routed gate/up on the other 33 MoE layers | `IQ2_XXS`, importance-matrix calibrated | Applies the most aggressive compression to the largest tensor bank. | | Routed down on the other 33 MoE layers | `Q2_K`, importance-matrix calibrated | More conservative two-bit storage on the projection that writes expert output to the residual stream. | | Attention projections | `Q8_0` | Protects a dense path used for every token. | | Shared experts | `Q8_0` | Protects the expert path active for every token. | | Vocabulary/output head | `Q8_0` | Protects final-logit fidelity. | | Indexer `attn_q_b` tensors on layers 2, 4, …, 42 | forced `F16` | Preserves a small, sensitive compressed-attention component. | | Remaining control, normalization, routing and auxiliary tensors | template-declared `F16`/`F32`/`I32` | Avoids forcing small or numerically sensitive tensors into the low-bit expert rules. | Observed main-GGUF inventory from strict DS4 inspection: | GGUF type | Tensors | | --- | ---: | | `F32` | 492 | | `F16` | 359 | | `I32` | 3 | | `Q8_0` | 345 | | `IQ2_XXS` | 66 | | `Q2_K` | 33 | | `MXFP4` | 30 | | **Total** | **1,328** | The target GGUF is version 3, describes approximately 284.33 billion logical parameters and retains the checkpoint's declared 1,048,576-token training context. That declaration is not evidence that a one-million-token inference run fits or remains robust on a particular machine. ### DSpark support model This is the actual 0731 three-stage DSpark module (`mtp.0`–`mtp.2`), not the legacy single-stage MTP attachment. The target model remains authoritative and verifies speculative proposals. | Parameter | Value | | --- | --- | | Stages | 3 | | Proposal block size | 5 | | Target layers | 40, 41, 42 | | Markov rank | 256 | | Noise token ID | 128799 | | Routed gate/up | `IQ2_XXS` (6 tensors) | | Routed down | exact native `MXFP4` (3 tensors) | | Dense projections | `Q8_0` (31 tensors) | | Control/auxiliary tensors | 7 `F16` + 34 `F32` tensors | | Total | 81 tensors | The support GGUF describes approximately 19.85 billion logical parameters and must be loaded alongside its matching target GGUF. ## Importance calibration The routed-expert calibration source is the 0731-native matrix from [`ox-ox/DeepSeek-V4-Flash-0731-GGUF`](https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-GGUF): - Repository revision: `6d58a3a36030c3ccb969bb5759fc6ae08cd299f8` - File: `imatrix/DeepSeek-V4-Flash-0731-chat-v2-routed-moe-ds4-1p5m.dat` - Size: `450,892,654` bytes - SHA-256: `6fce7674df701de544e5d3351aab04e67602eddeafeb48cf70e77ebe47239eb4` - Coverage: exactly 129 routed target tensors (43 layers × gate/up/down), with exact name and vector-dimension validation The build ran strict imatrix validation. Preserved native MXFP4 tensors do not consume imatrix data because they are copied from the official FP8 checkpoint's packed source representation rather than requantized. The published matrix has no native DSpark entries. The support-model importer therefore made deterministic target-layer proxy aliases: - `mtp.0` ← target layer 40 - `mtp.1` ← target layer 41 - `mtp.2` ← target layer 42 The extended matrix contains 138 entries and has SHA-256 `689b446ed2e2657ebcb69a6516781ee0444e07641aeb6125233fb4e6fe7cbce3`. This is an explicitly recorded proxy, not a fresh activation capture from the DSpark drafter. ## Weight provenance and metadata template The sole source of weight values was the local copy of `deepseek-ai/DeepSeek-V4-Flash-0731` at official revision `9e165c30e2704aec5d9d593cce3eebd58bbef1cb`. All 48 FP8 weight shards— `166,886,535,336` bytes in total—were fully SHA-256 verified and guarded against mutation throughout conversion. The exact-0731 GGUF from [`antirez/deepseek-v4-gguf`](https://huggingface.co/antirez/deepseek-v4-gguf) was used only as a bounded metadata, tokenizer, tensor-order and shape template: - Repository revision: `1cd7b564460821938add0475a60b942c409295e0` - Template LFS SHA-256: `ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0` - Template Xet object: `7da16e1025c1b856490c29c341f4e467d15cb195389c70383dade5e6108799ac` - Verified 64 MiB header SHA-256: `f0e1d5e8f3b008402aa6eb32cada3873dd926c8bf5e7d00d7788eec65f09dd6d` - Template weight values used: **no** ## Runtime requirement Use [a recent ds4 version from the main branch](https://github.com/antirez/ds4). This is required because the main GGUF contains 30 native MXFP4 routed-expert tensors. An older runtime without the MXFP4 loader and Metal kernels now in `main` is not compatible, even if it can parse the GGUF header. Install or update a `main`-branch checkout: ```bash git clone https://github.com/antirez/ds4.git cd ds4 git switch main git pull --ff-only make ``` Build note: the [`d516d4e` quantizer PR](https://github.com/apetersson/ds4-omlx/tree/d516d4eeb82c454aeb2831af1b1961801d6b571b), tracked as [`antirez/ds4#642`](https://github.com/antirez/ds4/issues/642), was needed to create the preserved-MXFP4 DSpark GGUF. Users do **not** need that PR to run either supplied file. Do not substitute `antirez/ds4` `main`, an older DS4 binary or a generic GGUF runtime. Container parsing alone does not demonstrate correct native MXFP4 or mixed DSpark execution. ## Running with DS4 Target-only Metal inference: ```bash ./ds4 --metal \ -m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf ``` Greedy DSpark inference: ```bash ./ds4 --metal \ -m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \ --mtp DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf \ --dspark --temp 0 ``` DSpark is opt-in, accelerates generation rather than prefill, and can be neutral or slower when proposal acceptance is low. Sampled decoding does not use DSpark proposals in the pinned runtime. Structural inspection: ```bash ./ds4 --cpu --inspect \ -m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \ --mtp DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf \ --dspark-strict ``` ## Memory planning on a 128 GB Mac The target machine has 128 GiB (`137,438,953,472` bytes) of unified physical memory, but the more relevant ceiling for a large Metal allocation is its approximately `121.60 GiB` `recommendedMaxWorkingSetSize`. The remaining roughly `6.40 GiB` is not extra model capacity; macOS, applications, drivers and untracked transient allocations still need memory. A configuration at or below 128 GiB can therefore be unsafe when it is above—or too close to—the Metal working-set recommendation. ### Fixed weight cost | Loaded weights | GiB | Share of 128 GiB | Share of 121.60 GiB Metal recommendation | | --- | ---: | ---: | ---: | | Target model only | `95.7644` | `74.8%` | `78.8%` | | Target + DSpark support | `102.5609` | `80.1%` | `84.3%` | These are GGUF payload sizes, before KV cache, indexed-attention scratch, prefill workspace, verifier state and other runtime allocations. DS4's planner uses an approximately `97.63 GiB` resident span for the main model after its mapping/alignment accounting, rather than treating the main file size as the entire live allocation. ### Context-length scaling at prefill chunk 1,024 `--ctx` is the total prompt-plus-completion capacity. `--prefill-chunk` is the maximum prompt microbatch DS4 processes at once; it is the relevant “batch size” for this single-session memory calculation. A larger chunk can improve prefill throughput, but indexed-attention scratch grows with both context length and chunk size. The following are conservative planning estimates. “Margin” is remaining space under the `121.60 GiB` Metal recommendation, not free system RAM. | Context | DSpark off: total | Off: margin | DSpark on: total | On: margin | | ---: | ---: | ---: | ---: | ---: | | 4,096 | `97.77 GiB` | `23.83 GiB` | `104.57 GiB` | `17.03 GiB` | | 32,768 | `98.05 GiB` | `23.55 GiB` | `104.85 GiB` | `16.75 GiB` | | 131,072 | `98.99 GiB` | `22.61 GiB` | `105.79 GiB` | `15.81 GiB` | | 262,144 | `100.24 GiB` | `21.36 GiB` | `107.04 GiB` | `14.56 GiB` | | 524,288 | `102.75 GiB` | `18.85 GiB` | `109.55 GiB` | `12.05 GiB` | | 1,048,576 | `107.77 GiB` | `13.83 GiB` | `114.56 GiB` | `7.04 GiB` | At ordinary 4K–128K contexts, the model weights dominate and chunk 1,024 leaves substantial planned margin. At 512K and especially 1M, the compressed KV and context-by-chunk attention workspace become material. DSpark adds approximately `6.80 GiB` at every context length because its support model remains resident during prefill even though speculative decoding only accelerates generation. ### Prefill batch-size effect at one-million-token context This table holds `--ctx 1048576` constant and varies `--prefill-chunk`. Totals are the conservative envelope of DS4's resident-span/context estimator and the completed release planning model. Percentages use all 128 GiB of physical RAM; the Metal margin remains the safer operational measure. | Prefill chunk | DSpark off total | Off: 128 GiB used | Off: Metal margin | DSpark on total | On: 128 GiB used | On: Metal margin | | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | 256 | `106.20 GiB` | `83.0%` | `15.40 GiB` | `113.00 GiB` | `88.3%` | `8.60 GiB` | | 512 | `106.72 GiB` | `83.4%` | `14.88 GiB` | `113.52 GiB` | `88.7%` | `8.08 GiB` | | 1,024 | `107.77 GiB` | `84.2%` | `13.83 GiB` | `114.56 GiB` | `89.5%` | `7.04 GiB` | | 2,048 | `110.30 GiB` | `86.2%` | `11.30 GiB` | `117.09 GiB` | `91.5%` | `4.51 GiB` | | 4,096 | `116.44 GiB` | `91.0%` | `5.16 GiB` | `123.24 GiB` | `96.3%` | **`−1.64 GiB`** | The 4,096/DSpark combination exceeds Metal's recommendation despite fitting numerically inside 128 GiB and should not be treated as resident-safe. The 2,048/DSpark combination is also tight: its `4.51 GiB` planned Metal margin can be consumed by omitted driver and verifier peaks. Chunk 1,024 is the sensible first DSpark experiment at 1M; target-only chunk 2,048 is the more conservative maximum-quality starting point. Chunks 256–512 provide more margin at the cost of more prefill iterations and likely lower prompt-processing throughput. Target-only 1M starting point: ```bash ./ds4 --metal \ -m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \ --ctx 1048576 --prefill-chunk 2048 ``` DSpark 1M starting point, only after unloading other large applications and measuring the target-only peak: ```bash ./ds4 --metal \ -m DeepSeek-V4-Flash-0731-DS4-Quality128.gguf \ --ctx 1048576 --prefill-chunk 1024 \ --mtp DeepSeek-V4-Flash-0731-DS4-Quality128-DSpark-support.gguf \ --dspark --temp 0 ``` ### Multiple sessions and server batching The tables describe one resident session. They do not mean that `N` concurrent requests cost exactly `N` times the displayed total: weights and some graph workspace are shared, while each resident session needs its own KV/context buffers and live request state. DS4 reports the aggregate minimum context-buffer request at server startup. For multi-user serving, start with one session, measure the process and system peak, then raise concurrency one session at a time; do not infer a safe concurrency count from the GGUF sizes alone. All figures above are planning estimates, not measurements of this exact artifact. Metal load, peak-memory capture and an actual 1,048,576-token run remain required. Keep several GiB of additional operational margin, watch memory pressure rather than only Activity Monitor's process RSS, and expect other active models or large applications to invalidate the table. ## Reproducibility - Build tooling revision: `e24746463a0a3e79036dd9c0472deac6ce704f08` - Build driver SHA-256: `cb1569b909de010b4aaa2a1b1b7ed8e277bcdf41e32585f1651005b41e81567f` - Build-time DS4/quantizer revision: `d516d4eeb82c454aeb2831af1b1961801d6b571b` - oMLX revision: `76352ed2363e42b2146243463875756e683a640d` - Profile SHA-256: `bc7f6470a0cb6208addf2ac52806d976108517119b411023da3fd91b492b218c` - Build run ID: `20260801-161944-54912` The tooling worktree was intentionally dirty and is cryptographically described in `BUILD_MANIFEST.json`, alongside the complete profile, source-shard hashes, calibration/template identities and exact converter commands. ## Validation status Completed on 2026-08-01: - focused build-driver and finalization test suites: 40/40 passed; - every one of the 48 official source weight shards fully SHA-256 verified; - strict routed-imatrix name and vector-dimension coverage passed; - main GGUF: exact size, 1,328 tensors, exact name set and exact type histogram; - DSpark GGUF: exact size, 81 tensors, three stages and target layers 40/41/42; - strict DSpark binding: 81 tensors, 0 missing, 0 invalid and 0 metadata errors; - `ds4 --cpu --inspect --dspark-strict` passed; - byte-for-byte regeneration passed for all 30 main and all three DSpark native MXFP4 tensors; and - the original build payloads passed their recorded SHA-256 checks before atomic publication; this README was added afterward and independently added to `SHA256SUMS`. Still required before making runtime, quality or performance claims: - successful Metal load and deterministic generation with DSpark disabled; - successful Metal generation with DSpark enabled and lossless target agreement; - proposal acceptance rate and accepted tokens per target step; - prompt-processing and generation tokens/second; - measured peak unified memory and practical context limits on the target host; - an actual 1,048,576-token context run; and - capability/perplexity comparisons against the official FP8 source. ## Limitations and responsible use Ultra-low-bit expert quantization can reduce reasoning, factuality, style fidelity and long-context robustness even when dense paths and selected experts are protected. Native MXFP4 Metal kernels and the mixed DSpark path are newer than DS4's mature Q4_K path. This package does not guarantee correctness, neutrality, safety, regulatory compliance or a particular response style. Evaluate it for the intended workload and apply appropriate access controls. ## License and attribution The upstream repository and weights are MIT licensed. This quantized derivative retains that license. Credit DeepSeek-AI for the original model, ox-ox for the 0731 routed-expert importance matrix, antirez for DS4 and the exact-0731 metadata recipe, and apetersson for the DS4 fork, conversion profile and release tooling.