Instructions to use 0xSero/DeepSeek-V4-Flash-0731-EXL3-2.5bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Trellis
How to use 0xSero/DeepSeek-V4-Flash-0731-EXL3-2.5bpw with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
DeepSeek-V4-Flash-0731-EXL3-2.5bpw
Campaign status: pending. This public model card reserves the release destination while the pinned calibration, FP4-to-EXL3 transcode, checksums, runtime validation, and upload run. No model weights are present yet.
Planned routed-expert EXL3/MCG transcode of deepseek-ai/DeepSeek-V4-Flash-0731 at 2.5 bpw. The source routed experts are packed E2M1 FP4 with UE8M0 scales, so this is an FP4-to-EXL3 transcode, not a BF16-to-EXL3 quantization.
Release gates
- Source revision:
9e165c30e2704aec5d9d593cce3eebd58bbef1cb - One sample-isolated 1,048,576-token natural-routing capture
- Full 43-layer × 256-expert route-coverage report
- Checkpoint-derived valid-token coverage for all three hash-routing layers, capped at 25% of calibration with no forced expert activation
- Exact clamped-SwiGLU down-projection calibration
- Complete allocation manifest, tensor inventory, SHA-256 checksums, and sizes
- Hybrid runtime dispatch: EXL3 backbone experts plus carried DeepSeek FP8 attention/dense/shared/MTP tensors
- TP4 CUDA-graph runtime validation with expert parallelism and MTP disabled
- Coding smoke and Terminal-Bench 2.1 reported only if actually measured
- K2/K2.5 fast runtime: blocked on a CUDA-graph-compatible fused kernel.
This is not planned as a stock Transformers or stock ExLlamaV3 checkpoint. The completed card will pin the required DeepSeek-V4-aware vLLM runtime and will state unsupported paths explicitly. No eager-mode fallback will be used.
All variants
| Target | Repository |
|---|---|
| 2.0 bpw | 0xSero/DeepSeek-V4-Flash-0731-EXL3-2.0bpw |
| 2.5 bpw | 0xSero/DeepSeek-V4-Flash-0731-EXL3-2.5bpw |
| 3.0 bpw | 0xSero/DeepSeek-V4-Flash-0731-EXL3-3.0bpw |
| 3.5 bpw | 0xSero/DeepSeek-V4-Flash-0731-EXL3-3.5bpw |
| 4.0 bpw | 0xSero/DeepSeek-V4-Flash-0731-EXL3-4.0bpw |
| 4.5 bpw | 0xSero/DeepSeek-V4-Flash-0731-EXL3-4.5bpw |
| 5.0 bpw | 0xSero/DeepSeek-V4-Flash-0731-EXL3-5.0bpw |
Suite index: 0xSero/DeepSeek-V4-Flash-0731-EXL3
Credits
Thanks to DeepSeek for the source model, TurboDerp for ExLlamaV3 and EXL3, and JarvisLabs for the planned campaign compute. This is an independent community transcode, not an official DeepSeek or TurboDerp release.
Model tree for 0xSero/DeepSeek-V4-Flash-0731-EXL3-2.5bpw
Base model
deepseek-ai/DeepSeek-V4-Flash-0731