DeepSeek-V4-Flash-W4A16-FP8-MTP — Ada (SM89) Serving Edition
This repository is a byte-identical mirror of canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP — all credit for the quantization and the checkpoint itself goes to canada-quant. The mirror exists so that the complete, verified recipe for serving this model on SM89 (Ada) GPUs — RTX 4090, RTX 6000 Ada — can live next to the weights: the setup script, the one required source patch, a systemd unit, and this document.
The Ada adaptation — four nights of debugging across two CUDA toolchains, three attention backends, and one 870-commit upstream regression — was worked out with the help of Claude (Anthropic).
What this checkpoint is
DeepSeek-V4-Flash quantized to W4A16 INT4 experts + FP8 attention (compressed-tensors) with a BF16 MTP head for speculative decoding. ~159 GB on disk, DeepSeek sparse attention (Lightning Indexer / MLA), 262,144-token context. It loads at roughly 38 GiB per GPU under tensor-parallel 4, which is exactly why 4× 48 GB Ada cards are such an attractive target — and exactly why it is painful that the default serving stacks refuse to run there.
262k is the checkpoint's context ceiling, not the recommended serving length on 48 GB Ada. After weights + activations + CUDA-graph pools, only ~1 GiB/GPU is left for KV cache, so a full 262k context fits only at high
--gpu-memory-utilizationand leaves almost no headroom. For concurrent or agentic workloads the generated defaults cap--max-model-lenlower (98304) to keep KV and JIT headroom — see Serving on 48 GB Ada below.
Measured performance
4× RTX 6000 Ada (48 GB, PCIe-only, TP=4), fp8_ds_mla KV cache,
262,144 max context, measured 2026-07-21:
| Configuration | Generation throughput | Notes |
|---|---|---|
| Plain decoding | ~80 tok/s single-stream | Triton sparse-MLA path |
MTP, num_speculative_tokens=1 |
115–133 tok/s sustained | draft acceptance 96.6–100%, mean acceptance length ~1.97–2.00 / 2.00 |
The BF16 MTP head drafts nearly perfectly on this checkpoint; a ~1.5–1.65× speedup from a single speculative token is about as good as MTP-k=1 gets.
These figures are single-stream throughput at the maximum-context operating
point (262k, high utilization, --max-num-batched-tokens 2048). That is the
right point to quote peak tok/s, but not the right point to run a
concurrent or agentic workload on 48 GB cards — see the next section for why,
and for the defaults the setup script actually generates.
Why Ada needs special handling
The stock DeepSeek-V4 serving paths gate on newer architectures at three independent layers: FlashMLA's sparse kernels (SM90+), DeepGEMM's FP8 paged-MQA-logits kernels (arch 9/10/12 only — no SM89 device code exists), and FlashInfer's sparse-MLA route. The path that actually works on SM89 is the portable Triton sparse-MLA implementation (upstream vLLM PR #41834) extended to Ada in a community fork.
The exact working stack
Everything below is automated by
setup_deepseek_v4_sm89.sh in this repo.
- vLLM fork: yhfgyyf/vllm-deepseek-v4-sm89,
release
v0.23.1rc1.dev145-g8c631d45e-cu130-sm89(July 1). The newerdev1018release does not work on SM89 — an upstream sync regressed the fork's SM89 indexer-logits fallbacks (sm12x_deep_gemm_fallbacks), so its decode path dead-ends in DeepGEMM's architecture assert. Install thedev145wheel specifically. - FlashInfer:
flashinfer_python 0.6.14+sm89(shipped as an asset of the fork'sdev1018release) +flashinfer-cubin==0.6.13. Idle on the Triton path, but part of the validated environment. - torch 2.11.0+cu130, with the CUDA 13.0 toolkit pinned at runtime
(
CUDA_HOME/CUDA_PATH/PATH). Mixed toolchains (an apt CUDA 12nvccshadowing 13.x headers) break JIT builds with walls of one-argument__cudaLauncherrors. - transformers pinned to
5.8.1, installed last. Newer releases reject this checkpoint's config: theirALLOWED_LAYER_TYPESvalidation no longer accepts thehash_moeentries inmlp_layer_types. Installing the vLLM wheel afterwards would let pip drag a newer transformers back in, so the pin must be the final install step. - deep_gemm must NOT be installed. The fork binds DeepGEMM's mqa-logits
kernels on presence, not capability — if the package is importable, SM89
decode crashes in
csrc/apis/attention.hpp("Unsupported architecture"). Absent, the capability-gated Triton fallbacks take over. - Launch essentials:
VLLM_TRITON_MLA_SPARSE=1(the master switch for the SM89 Triton path — load-bearing), no--attention-backendflag (forcing the FlashInfer backend crashes on an SWA-cache shape mismatch),--kv-cache-dtype fp8_ds_mla --block-size 256, and thedeepseek_v4tokenizer / reasoning-parser / tool-call-parser modes. - MTP:
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'.--max-num-batched-tokensis2048at the peak-throughput point above; the generatedserve.shraises it to4096for the agentic-serving defaults (which enables chunked prefill — see the memory section, as its buffers come out of the KV pool).
The one required source patch (MTP only)
The fork's fused o-projection (nvidia/ops/o_proj.py,
deep_gemm_fp8_o_proj) unconditionally reads FP8 block scales
(weight_scale / weight_scale_inv) off wo_a. The FP8 target model sails
through; the checkpoint's BF16 MTP drafter head has no scales, so
enabling speculative decoding crashes with
AttributeError: 'ColumnParallelLinear' object has no attribute 'weight_scale'.
patch_dev145_bf16_oproj.py inserts a
guarded fallback that fires only for scale-less modules: inverse RoPE
replicated in plain torch (interleaved even/odd pairs on the last rope_dim
dims of each head, rotation by −θ, mirroring the fork's own
_fused_inv_rope_fp8_quant_per_head kernel), a grouped BF16 einsum against
wo_a.weight viewed [n_groups, o_lora_rank, heads_per_group·head_dim],
then wo_b. The FP8 path is untouched down to the byte, and the drafter is
a single tiny layer, so the unfused fallback costs nothing measurable. The
patcher is idempotent, backs up the original, and refuses to apply against
any file that doesn't match the dev145 layout.
Serving on 48 GB Ada: memory, concurrency, and context
On these cards KV cache is the thin slice — weights, activations, and CUDA-graph pools consume almost the whole 48 GB, leaving on the order of ~1 GiB/GPU for KV. Two failure modes sit on opposite sides of a narrow tuning window, and the generated defaults are the corner that clears both:
- JIT-OOM (too little headroom). The Triton sparse-MLA path compiles a
kernel specialization per prompt shape, lazily, and loading each freshly
compiled cubin (
cuModuleLoadData) needs driver-level free VRAM outside the torch pool. At high utilization a never-before-seen prompt shape arriving mid-run can find no room to load its kernel and crash a worker (observed in_tf32_hc_prenorm_gemm), taking the server down. Two things defuse this: the generatedwarmup.shcompiles the shape space at startup while memory is free, and the systemd unit usesRestart=always(a JIT-OOM can exit cleanly from systemd's view, soon-failureis not enough). - KV-fit refusal (too little KV). vLLM refuses to boot if a single full
--max-model-lensequence won't fit the KV pool. At the headroom-safe utilizations below, full 262k does not fit;--max-model-len 98304does, with room for the KV a long agentic trajectory actually uses.
Concurrency: use --max-num-seqs 1 for agentic workloads. With several
30–100k-token trajectories in flight, they cannot co-reside in the KV pool, so
vLLM preempts and re-prefills evicted sequences from scratch — prefix-cache
hit rate collapses below 40%, generation throughput drops to single digits
while prefill spikes, and wall-clock blows up roughly 10×. A single resident
stream (prefix hit ~85–91%) decoding continuously is dramatically faster in
practice. One stream that never evicts beats four that thrash.
Utilization. 0.97 maximizes the KV pool and works for a watched run
backed by warmup + auto-restart, but leaves only ~0.5–0.8 GiB/GPU free under a
long trajectory. For unattended or shared operation prefer
--gpu-memory-utilization 0.95: it restores ~2.4 GiB of headroom at identical
seq=1 throughput. The setup script exposes both as parameters; pick per how
closely you'll be watching.
Smaller Ada cards (RTX 4090, 24 GB). This recipe is validated on 4× 48 GB
(RTX 6000 Ada). A 159 GB checkpoint under TP=4 needs ~38 GiB/GPU for weights
alone, so it does not fit four 24 GB 4090s — you would need more cards
(e.g. TP=8) and an even tighter KV budget, and none of the memory figures here
transfer directly. The rtx-4090 tag reflects the SM89 architecture the stack
targets, not a claim that this checkpoint fits on 4090-class VRAM.
Quickstart
# 1. Build the environment (venv, pinned wheels, patch, verification,
# generated serve.sh / warmup.sh / canary.sh / systemd unit):
./setup_deepseek_v4_sm89.sh ~/deepseek-v4-serve SinclairSchneider/DeepSeek-V4-Flash-W4A16-FP8-MTP-Ada
# (the original repo id works identically:
# ./setup_deepseek_v4_sm89.sh ~/deepseek-v4-serve canada-quant/DeepSeek-V4-Flash-W4A16-FP8-MTP)
# 2. Fetch the weights once (~159 GB):
~/deepseek-v4-serve/.venv/bin/huggingface-cli download SinclairSchneider/DeepSeek-V4-Flash-W4A16-FP8-MTP-Ada
# 3. Serve (OpenAI-compatible API on :8002):
~/deepseek-v4-serve/serve.sh
# 4. Smoke test:
~/deepseek-v4-serve/canary.sh
Expect ~80 s of weight loading plus ~5 s for the MTP drafter, then a Triton
shape-warmup pass (warmup.sh, also generated) that compiles the kernel
specializations up front, then steady-state throughput. For unattended
operation the setup script generates a systemd unit
(see deepseek-v4-flash.service) that runs the
warmup automatically after every start (ExecStartPost) and uses
Restart=always so a mid-run kernel-load OOM self-heals rather than leaving a
dead port.
Gotchas the scripts already handle
- pip rejects wheels whose filenames were changed — release assets are kept
under their canonical names, and
+in GitHub asset URLs is sent as%2B. - Stale
VLLM_TRITON_MLA_SPARSE_*exports in shell profiles are silently consumed bydev145and can corrupt cache geometry —serve.shscrubs them withenv -u, and systemd's clean service environment makes the unit immune by construction. HF_HUB_OFFLINE=1resolution followsrefs/main; an interrupted online metadata refresh can leave that ref pointing at a snapshot that was never downloaded.serve.shtherefore resolves the model to a concrete snapshot directory itself, trusting the ref only when its target exists on disk.
Credits
- canada-quant — the original DeepSeek-V4-Flash-W4A16-FP8-MTP quantization this repository mirrors. All credit for the checkpoint.
- yhfgyyf — the
vllm-deepseek-v4-sm89
fork whose
dev145release contains the working SM89 path, building on vLLM's portable Triton sparse-MLA kernels (PR #41834). - DeepSeek — the DeepSeek-V4-Flash model family.
- Claude (Anthropic) — the Ada adaptation, root-cause analysis, the BF16 o_proj patch, and these scripts were developed with Claude's help.
The weights inherit the license of the original repository.
- Downloads last month
- 89
Model tree for SinclairSchneider/DeepSeek-V4-Flash-W4A16-FP8-MTP-Ada
Base model
deepseek-ai/DeepSeek-V4-Flash