File size: 1,113 Bytes
01b5998
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
---
language: en
license: mit
pipeline_tag: text-generation
tags:
- mlx
library_name: mlx
base_model: deepreinforce-ai/Ornith-1.0-35B
---

This model was converted to MLX format and quantized from [Ornith-1.0-35B](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B) using [oMLX](https://github.com/jundot/omlx).

>[!CAUTION]
>The conversion required to stack expert MLP weights into fused per-layer tensors. Treat this as an experiment and act accordingly.

## What is "oQ"?

See ["oQ: oMLX Universal Dynamic Quantization"](https://github.com/jundot/omlx/blob/main/docs/oQ_Quantization.md) for details.

## What is "VL"?

"VL" is Vision-Language, meaning quantization preserves the original model's multimodality.

No "VL" means quantization is Text-Only.

## What is "FP16"?

"FP16" is an M1/M2 Apple Silicon tweak that delivers a very noticeable prompt processing boost, because older M-series lack native BF16 hardware support. See ["Metal FP32 Vs BF16 Vs FP16 benchmark"](https://github.com/deepsweet/metal-fp32-bf16-fp16) for details.

No "FP16" means quantization is better suited for M3+ Apple Silicon.