--- language: en license: mit pipeline_tag: text-generation tags: - mlx library_name: mlx base_model: deepreinforce-ai/Ornith-1.0-35B --- This model was converted to MLX format and quantized from [Ornith-1.0-35B](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B) using [oMLX](https://github.com/jundot/omlx). >[!CAUTION] >The conversion required to stack expert MLP weights into fused per-layer tensors. Treat this as an experiment and act accordingly. ## What is "oQ"? See ["oQ: oMLX Universal Dynamic Quantization"](https://github.com/jundot/omlx/blob/main/docs/oQ_Quantization.md) for details. ## What is "VL"? "VL" is Vision-Language, meaning quantization preserves the original model's multimodality. No "VL" means quantization is Text-Only. ## What is "FP16"? "FP16" is an M1/M2 Apple Silicon tweak that delivers a very noticeable prompt processing boost, because older M-series lack native BF16 hardware support. See ["Metal FP32 Vs BF16 Vs FP16 benchmark"](https://github.com/deepsweet/metal-fp32-bf16-fp16) for details. No "FP16" means quantization is better suited for M3+ Apple Silicon.