pipenetwork/DeepSeek-V4-Flash-MLX-8bit
Text Generation • 81B • Updated • 27
304B MoE on Apple Silicon. deepseek_v4 is in no released runtime — ported from scratch. Take mixed-4_8bit. github.com/PipeNetwork/deepseek-v4-mlx
Note 304 GB · ppl 5.9878 · highest fidelity
Note 233 GB · ppl 6.0260 · +0.6% vs 8-bit for 70 GB less
Note RECOMMENDED · 165 GB · ppl 6.1262 · 4-bit routed experts, 8-bit everything else. Experts are 98.1% of params, so raising the other 2% costs 3 GB and recovers 56% of the gap to 8-bit.
Note 162 GB · ppl 6.3005 · superseded by mixed-4_8bit — 3 GB more, 2.8% better
Note 129 GB · ppl 6.4025 · 25% of experts pruned; statistically indistinguishable from 4-bit at 33 GB less
Note 111 GB · ppl 6.9634 · 37% pruned; fits a 128 GB Mac
Note 93 GB · ppl 7.7520 · 50% pruned; smallest build