| --- |
| |
| pipeline_tag: text-generation |
| base_model: canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP |
| tags: |
| - deepseek |
| - deepseek-v4 |
| - w4a16 |
| - fp8 |
| - mtp |
| - speculative-decoding |
| - sm89 |
| - ada |
| - rtx-4090 |
| - rtx-6000-ada |
| - vllm |
| --- |
| |
| # DeepSeek-V4-Flash-W4A16-FP8-MTP — Ada (SM89) Serving Edition |
|
|
| This repository is a **byte-identical mirror** of |
| [canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP](https://huggingface.co/canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP) |
| — all credit for the quantization and the checkpoint itself goes to |
| **canada-quant**. The mirror exists so that the complete, verified recipe for |
| serving this model on **SM89 (Ada) GPUs** — RTX 4090, RTX 6000 Ada — can live |
| next to the weights: the setup script, the one required source patch, a |
| systemd unit, and this document. |
|
|
| The Ada adaptation — four nights of debugging across two CUDA toolchains, |
| three attention backends, and one 870-commit upstream regression — was worked |
| out with the help of **Claude (Anthropic)**. |
|
|
| ## What this checkpoint is |
|
|
| DeepSeek-V4-Flash quantized to **W4A16 INT4 experts + FP8 attention** |
| (compressed-tensors) with a **BF16 MTP head** for speculative decoding. |
| ~159 GB on disk, DeepSeek sparse attention (Lightning Indexer / MLA), |
| **262,144-token context**. It loads at roughly 38 GiB per GPU under |
| tensor-parallel 4, which is exactly why 4× 48 GB Ada cards are such an |
| attractive target — and exactly why it is painful that the default serving |
| stacks refuse to run there. |
|
|
| > **262k is the checkpoint's context ceiling, not the recommended serving |
| > length on 48 GB Ada.** After weights + activations + CUDA-graph pools, only |
| > ~1 GiB/GPU is left for KV cache, so a full 262k context fits only at high |
| > `--gpu-memory-utilization` and leaves almost no headroom. For concurrent or |
| > agentic workloads the generated defaults cap `--max-model-len` lower (98304) |
| > to keep KV *and* JIT headroom — see [Serving on 48 GB Ada](#serving-on-48-gb-ada-memory-concurrency-and-context) below. |
|
|
| ## Measured performance |
|
|
| 4× RTX 6000 Ada (48 GB, PCIe-only, TP=4), `fp8_ds_mla` KV cache, |
| 262,144 max context, measured 2026-07-21: |
|
|
| | Configuration | Generation throughput | Notes | |
| |---|---|---| |
| | Plain decoding | ~80 tok/s single-stream | Triton sparse-MLA path | |
| | **MTP, `num_speculative_tokens=1`** | **115–133 tok/s sustained** | draft acceptance 96.6–100%, mean acceptance length ~1.97–2.00 / 2.00 | |
|
|
| The BF16 MTP head drafts nearly perfectly on this checkpoint; a ~1.5–1.65× |
| speedup from a single speculative token is about as good as MTP-k=1 gets. |
|
|
| These figures are **single-stream throughput at the maximum-context operating |
| point** (262k, high utilization, `--max-num-batched-tokens 2048`). That is the |
| right point to *quote* peak tok/s, but not the right point to *run* a |
| concurrent or agentic workload on 48 GB cards — see the next section for why, |
| and for the defaults the setup script actually generates. |
|
|
| ## Why Ada needs special handling |
|
|
| The stock DeepSeek-V4 serving paths gate on newer architectures at three |
| independent layers: FlashMLA's sparse kernels (SM90+), DeepGEMM's FP8 |
| paged-MQA-logits kernels (arch 9/10/12 only — no SM89 device code exists), |
| and FlashInfer's sparse-MLA route. The path that actually works on SM89 is |
| the **portable Triton sparse-MLA implementation** (upstream vLLM |
| [PR #41834](https://github.com/vllm-project/vllm/pull/41834)) extended to |
| Ada in a community fork. |
|
|
| ## The exact working stack |
|
|
| Everything below is automated by |
| [`setup_deepseek_v4_sm89.sh`](./setup_deepseek_v4_sm89.sh) in this repo. |
|
|
| - **vLLM fork**: [yhfgyyf/vllm-deepseek-v4-sm89](https://github.com/yhfgyyf/vllm-deepseek-v4-sm89), |
| release |
| [`v0.23.1rc1.dev145-g8c631d45e-cu130-sm89`](https://github.com/yhfgyyf/vllm-deepseek-v4-sm89/releases/tag/v0.23.1rc1.dev145-g8c631d45e-cu130-sm89) |
| (July 1). **The newer `dev1018` release does not work on SM89** — an |
| upstream sync regressed the fork's SM89 indexer-logits fallbacks |
| (`sm12x_deep_gemm_fallbacks`), so its decode path dead-ends in DeepGEMM's |
| architecture assert. Install the `dev145` wheel specifically. |
| - **FlashInfer**: `flashinfer_python 0.6.14+sm89` (shipped as an asset of the |
| fork's `dev1018` release) + `flashinfer-cubin==0.6.13`. Idle on the Triton |
| path, but part of the validated environment. |
| - **torch 2.11.0+cu130**, with the **CUDA 13.0 toolkit pinned** at runtime |
| (`CUDA_HOME`/`CUDA_PATH`/`PATH`). Mixed toolchains (an apt CUDA 12 `nvcc` |
| shadowing 13.x headers) break JIT builds with walls of one-argument |
| `__cudaLaunch` errors. |
| - **transformers pinned to `5.8.1`, installed last.** Newer releases reject |
| this checkpoint's config: their `ALLOWED_LAYER_TYPES` validation no longer |
| accepts the `hash_moe` entries in `mlp_layer_types`. Installing the vLLM |
| wheel afterwards would let pip drag a newer transformers back in, so the |
| pin must be the final install step. |
| - **deep_gemm must NOT be installed.** The fork binds DeepGEMM's mqa-logits |
| kernels on *presence*, not capability — if the package is importable, SM89 |
| decode crashes in `csrc/apis/attention.hpp` ("Unsupported architecture"). |
| Absent, the capability-gated Triton fallbacks take over. |
| - **Launch essentials**: `VLLM_TRITON_MLA_SPARSE=1` (the master switch for |
| the SM89 Triton path — load-bearing), **no `--attention-backend` flag** |
| (forcing the FlashInfer backend crashes on an SWA-cache shape mismatch), |
| `--kv-cache-dtype fp8_ds_mla --block-size 256`, and the `deepseek_v4` |
| tokenizer / reasoning-parser / tool-call-parser modes. |
| - **MTP**: `--speculative-config '{"method":"mtp","num_speculative_tokens":1}'`. |
| `--max-num-batched-tokens` is `2048` at the peak-throughput point above; the |
| generated `serve.sh` raises it to `4096` for the agentic-serving defaults |
| (which enables chunked prefill — see the memory section, as its buffers come |
| out of the KV pool). |
|
|
| ## The one required source patch (MTP only) |
|
|
| The fork's fused o-projection (`nvidia/ops/o_proj.py`, |
| `deep_gemm_fp8_o_proj`) unconditionally reads FP8 block scales |
| (`weight_scale` / `weight_scale_inv`) off `wo_a`. The FP8 target model sails |
| through; the checkpoint's **BF16 MTP drafter head** has no scales, so |
| enabling speculative decoding crashes with |
| `AttributeError: 'ColumnParallelLinear' object has no attribute 'weight_scale'`. |
|
|
| [`patch_dev145_bf16_oproj.py`](./patch_dev145_bf16_oproj.py) inserts a |
| guarded fallback that fires only for scale-less modules: inverse RoPE |
| replicated in plain torch (interleaved even/odd pairs on the last `rope_dim` |
| dims of each head, rotation by −θ, mirroring the fork's own |
| `_fused_inv_rope_fp8_quant_per_head` kernel), a grouped BF16 einsum against |
| `wo_a.weight` viewed `[n_groups, o_lora_rank, heads_per_group·head_dim]`, |
| then `wo_b`. The FP8 path is untouched down to the byte, and the drafter is |
| a single tiny layer, so the unfused fallback costs nothing measurable. The |
| patcher is idempotent, backs up the original, and refuses to apply against |
| any file that doesn't match the `dev145` layout. |
|
|
| ## Serving on 48 GB Ada: memory, concurrency, and context |
|
|
| On these cards **KV cache is the thin slice** — weights, activations, and |
| CUDA-graph pools consume almost the whole 48 GB, leaving on the order of |
| ~1 GiB/GPU for KV. Two failure modes sit on opposite sides of a narrow tuning |
| window, and the generated defaults are the corner that clears both: |
|
|
| - **JIT-OOM (too little headroom).** The Triton sparse-MLA path compiles a |
| kernel specialization *per prompt shape*, lazily, and loading each freshly |
| compiled cubin (`cuModuleLoadData`) needs driver-level free VRAM *outside* |
| the torch pool. At high utilization a never-before-seen prompt shape |
| arriving mid-run can find no room to load its kernel and crash a worker |
| (observed in `_tf32_hc_prenorm_gemm`), taking the server down. Two things |
| defuse this: the generated `warmup.sh` compiles the shape space at startup |
| while memory is free, and the systemd unit uses `Restart=always` (a JIT-OOM |
| can exit *cleanly* from systemd's view, so `on-failure` is not enough). |
| - **KV-fit refusal (too little KV).** vLLM refuses to boot if a single full |
| `--max-model-len` sequence won't fit the KV pool. At the headroom-safe |
| utilizations below, full 262k does **not** fit; `--max-model-len 98304` does, |
| with room for the KV a long agentic trajectory actually uses. |
|
|
| **Concurrency: use `--max-num-seqs 1` for agentic workloads.** With several |
| 30–100k-token trajectories in flight, they cannot co-reside in the KV pool, so |
| vLLM preempts and *re-prefills* evicted sequences from scratch — prefix-cache |
| hit rate collapses below 40%, generation throughput drops to single digits |
| while prefill spikes, and wall-clock blows up roughly 10×. A single resident |
| stream (prefix hit ~85–91%) decoding continuously is dramatically faster in |
| practice. One stream that never evicts beats four that thrash. |
|
|
| **Utilization.** `0.97` maximizes the KV pool and works for a *watched* run |
| backed by warmup + auto-restart, but leaves only ~0.5–0.8 GiB/GPU free under a |
| long trajectory. For **unattended or shared** operation prefer |
| `--gpu-memory-utilization 0.95`: it restores ~2.4 GiB of headroom at identical |
| seq=1 throughput. The setup script exposes both as parameters; pick per how |
| closely you'll be watching. |
|
|
| **Smaller Ada cards (RTX 4090, 24 GB).** This recipe is validated on 4× 48 GB |
| (RTX 6000 Ada). A 159 GB checkpoint under TP=4 needs ~38 GiB/GPU for weights |
| alone, so it does **not** fit four 24 GB 4090s — you would need more cards |
| (e.g. TP=8) and an even tighter KV budget, and none of the memory figures here |
| transfer directly. The `rtx-4090` tag reflects the SM89 architecture the stack |
| targets, not a claim that this checkpoint fits on 4090-class VRAM. |
|
|
| ## Quickstart |
|
|
| ```bash |
| # 1. Build the environment (venv, pinned wheels, patch, verification, |
| # generated serve.sh / warmup.sh / canary.sh / systemd unit): |
| ./setup_deepseek_v4_sm89.sh ~/deepseek-v4-serve SinclairSchneider/DeepSeek-V4-Flash-W4A16-FP8-MTP-Ada |
| # (the original repo id works identically: |
| # ./setup_deepseek_v4_sm89.sh ~/deepseek-v4-serve canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP) |
| |
| # 2. Fetch the weights once (~159 GB): |
| ~/deepseek-v4-serve/.venv/bin/huggingface-cli download SinclairSchneider/DeepSeek-V4-Flash-W4A16-FP8-MTP-Ada |
| |
| # 3. Serve (OpenAI-compatible API on :8002): |
| ~/deepseek-v4-serve/serve.sh |
| |
| # 4. Smoke test: |
| ~/deepseek-v4-serve/canary.sh |
| ``` |
|
|
| Expect ~80 s of weight loading plus ~5 s for the MTP drafter, then a Triton |
| shape-warmup pass (`warmup.sh`, also generated) that compiles the kernel |
| specializations up front, then steady-state throughput. For unattended |
| operation the setup script generates a systemd unit |
| (see [`deepseek-v4-flash.service`](./deepseek-v4-flash.service)) that runs the |
| warmup automatically after every start (`ExecStartPost`) and uses |
| `Restart=always` so a mid-run kernel-load OOM self-heals rather than leaving a |
| dead port. |
|
|
| ## Gotchas the scripts already handle |
|
|
| - pip rejects wheels whose filenames were changed — release assets are kept |
| under their canonical names, and `+` in GitHub asset URLs is sent as `%2B`. |
| - Stale `VLLM_TRITON_MLA_SPARSE_*` exports in shell profiles are silently |
| consumed by `dev145` and can corrupt cache geometry — `serve.sh` scrubs |
| them with `env -u`, and systemd's clean service environment makes the unit |
| immune by construction. |
| - `HF_HUB_OFFLINE=1` resolution follows `refs/main`; an interrupted online |
| metadata refresh can leave that ref pointing at a snapshot that was never |
| downloaded. `serve.sh` therefore resolves the model to a concrete snapshot |
| directory itself, trusting the ref only when its target exists on disk. |
|
|
| ## Credits |
|
|
| - **[canada-quant](https://huggingface.co/canada-quant)** — the original |
| [DeepSeek-V4-Flash-W4A16-FP8-MTP](https://huggingface.co/canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP) |
| quantization this repository mirrors. All credit for the checkpoint. |
| - **[yhfgyyf](https://github.com/yhfgyyf)** — the |
| [vllm-deepseek-v4-sm89](https://github.com/yhfgyyf/vllm-deepseek-v4-sm89) |
| fork whose `dev145` release contains the working SM89 path, building on |
| vLLM's portable Triton sparse-MLA kernels |
| ([PR #41834](https://github.com/vllm-project/vllm/pull/41834)). |
| - **DeepSeek** — the DeepSeek-V4-Flash model family. |
| - **[Claude](https://claude.com) (Anthropic)** — the Ada adaptation, |
| root-cause analysis, the BF16 o_proj patch, and these scripts were |
| developed with Claude's help. |
| |
| The weights inherit the license of the original repository. |
| |