Instructions to use AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP" --prompt "Once upon a time"
- Atomic Chat
AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP
An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head is preserved at BF16 in the checkpoint (or a bound sidecar when present).
Checkpoint Tier 1 certified (experimental) on
df-macstudio-m2at Hub revisione22b117aa812b29943b160bb0fbf0b962d0d3819. Safetensors fingerprints are unchanged; this metadata repair is not a new MTP acceleration certificate.
Model details
| Property | Value |
|---|---|
| Base model | deepseek-ai/DeepSeek-V4-Flash |
| Source revision | 60d8d70770c6776ff598c94bb586a859a38244f1 |
| Product family | deepseek-v4 |
| Source architecture | DeepseekV4ForCausalLM (mixture of experts (MoE)); text path optimized |
| Main-model parameters | 284.33B logical parameters |
| Quantizer | AXQuant 1.5.1 |
| Hub budget class | 2bit |
| AXQuant base precision class | 2bit-experimental |
| Planned storage-adjusted BPW | 3.4232 |
| Measured main-model BPW | 3.1329 |
| Measured total BPW, including MTP | 3.1605 |
| Safetensors weight size | 114.94 GB |
| Approximate complete download | 115.02 GB |
| Configured maximum context | 1,048,576 tokens; practical limits depend on unified memory |
| Primary MLX runtime | MLX-LM |
| AX Engine native execution | Direct runtime smoke passed with AX Engine 6.15.0; MTP remains direct fallback |
| MTP present | True |
| Vision present | False |
| Audio present | False |
This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.
Choosing an AXQ pack
AXQ names describe a storage-budget product class, not one uniform precision applied to every
tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative.
In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting
6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection
floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily
protected models. When that collapse happens, AutomatosX does not publish a separate
misleading 4bit sibling for that base.
| Sibling | Intended trade-off |
|---|---|
| This 2bit pack | Lowest-storage AXQ budget; check its exact BPW |
| 4bit sibling | Higher average precision near the 4-BPW budget |
See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.
Download
python -m pip install -U huggingface_hub
hf download AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP --local-dir ./AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP
Allow at least 115.02 GB of free disk space. Pin the resulting Hub commit in reproducible
deployments rather than relying indefinitely on main.
Run with MLX-LM
python -m pip install -U mlx-lm
mlx_lm.generate \
--model AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP \
--prompt "Explain mixed-precision quantization in three sentences." \
--max-tokens 128 \
--temp 0.0
MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime
metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore
does not establish MTP acceleration or vision-language quality. The artifact records MLX
0.32.0 and MLX-LM 0.31.3 from conversion.
AX Engine and DeepSeek MTP status
The checkpoint Tier 1 certificate
is bound to Hub revision e22b117aa812b29943b160bb0fbf0b962d0d3819; every Safetensors LFS
fingerprint is unchanged on this metadata-only revision. AX Engine 6.15.0 passed direct load,
chat, stream, and context-retrieval smoke tests on df-macstudio-m2 with
AX_ENGINE_2BIT_EXPERIMENTAL=1. That checkpoint result does not certify speculative decode.
The packaged mtp.safetensors is the native DeepSeek V4 nextn sidecar, not a Qwen
qwen3-next-mtp sidecar. AX Engine 7.1.5 recognizes this layout but keeps the product route on
direct fallback until a revision-bound Tier 2 MTP acceptance, exactness, and speed certificate
exists. Stock MLX-LM runs the backbone without activating the sidecar, and the oMLX/MTPLX Qwen
import workflow does not apply. The internal DeepSeek MTP certification-candidate switch is for
the formal harness, not normal serving.
Quantization layout
| Main-weight precision | Parameters | Share |
|---|---|---|
2bit |
278.11B | 95.59% |
4bit |
3.64B | 1.25% |
8bit |
529.53M | 0.18% |
bf16 |
8.67B | 2.98% |
- Quantization methods:
affine, bf16. - Group sizes used by quantized assignments:
32. - MTP sidecar: 1575 tensors, 6.61B parameters, 3.59 GB, BF16, F32, F8_E4M3, F8_E8M0, I8.
- Vision sidecar: not included.
- Optimization scope:
text-path. - Support tier:
convertible.
BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.
Evidence and validation status
| Check | Status |
|---|---|
| Planning evidence | architecture_prior |
| Calibration | none; the allocation is based on architecture priors |
| Quantizer execution | 33492/33492 recorded module conversions succeeded; 0 fallbacks |
| AX Engine direct runtime | Passed on df-macstudio-m2 with AX Engine 6.15.0 |
| Quality versus BF16 or uniform baselines | Not published; no quality-retention claim |
| MTP acceptance and speed | not measured; no MTP speedup claim |
| AX Engine kernel evidence | unmeasured |
| Vision-language quality | Not applicable (no vision tower in this package) |
| Speech-recognition quality | Not applicable |
| Long-context quality | 1,048,576-token capacity is config metadata, not a validated claim |
| Release certification | Checkpoint Tier 1 certified (experimental) at e22b117a; MTP Tier 2 not certified |
Intended use and limitations
Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.
No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.
Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.
MTP uses the DeepSeek V4
nextncontract. Product serving remains direct fallback until a revision-bound Tier 2 certificate exists.The configured context window can require substantially more memory as the KV cache grows.
Direct AX Engine runtime passed at the certificate revision; this package still has no in-repo native manifest, and MTP Tier 2 remains unverified.
Upstream capabilities, limitations, biases, and responsible-use guidance still apply.
Provenance and audit files
axquant_manifest.json: package identity, byte accounting, runtime contract, software versions, and file checksums.axquant_plan.json: per-tensor precision decisions and planning evidence.axquant_quantizer_execution.json: conversion coverage and fallback records.axquant_runtime.json: declared AX Engine and MLX compatibility metadata; runtime checks remain separate evidence.axquant_mtp_sidecar_manifest.json: MTP tensor provenance.
All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. If an OptiQ repository is published separately, it uses a different quantizer and should not be assumed to have identical BPW or quality.
License
The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the deepseek-ai/DeepSeek-V4-Flash model card for license terms, model limitations, and responsible-use guidance.
- Downloads last month
- 767
2-bit
Model tree for AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP
Base model
deepseek-ai/DeepSeek-V4-Flash