Instructions to use miabchdave/Qwen3.6-27B-oQ8e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use miabchdave/Qwen3.6-27B-oQ8e-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwen3.6-27B-oQ8e-mtp miabchdave/Qwen3.6-27B-oQ8e-mtp
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
Qwen3.6-27B-oQ8e-mtp
This model was quantized using oQe (oMLX v0.5.1) mixed-precision quantization directly from Qwen3.6 27B BF16. SHA256 shows this oQ8e Quantized Version is different than previously generated Qwen3.6 27B oQ8e versions on HF as of the first commit date so it's likely using a more current version of oQe.
MTP Head and Vision Tower is present. Model is MLX safetensors optimized for Apple Silicon performance.
Code generated at ~45 TG/S on M5 Max with 8K Context using oMLX v5.1 with Lightning MTP Option ON in oMLX.
Uses froggeric chat_template.jinja (https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates) to optimize thinking and correct other issues in the default chat_template. Return the chat_template.jinja to version in HF Qwen3.6 repo for any issues.
This quantization is based upon Qwen3.6 27B
Quantization details
- Bits: 8
- Group size: 64
- Format: MLX safetensors
- Downloads last month
- 252
8-bit