Ornith-1.0-35B-oQ2 / README.md
jasen215's picture
Upload README.md with huggingface_hub
4d27173 verified
|
Raw
History Blame Contribute Delete
2 kB
metadata
library_name: mlx
tags:
  - mlx
  - oq
  - quantized
  - benchmark
  - performance
  - moe
  - ornith

Ornith-1.0-35B-oQ2

This model was quantized using oQ (oMLX v0.5.3) mixed-precision quantization.

Base model: Ornith-1.0-35B

Quantization details

  • Model type: qwen3_5_moe
  • Bits: 2
  • Group size: 64
  • Format: MLX safetensors

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.5.3
  • Max Concurrent Requests: 4
  • Settings:
    • Thinking: Disabled
    • TurboQuant KV Cache: Enabled

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 1063.3 20.83 963.0 tok/s 48.4 tok/s 3.722 309.5 tok/s 12.62 GB
pp4096/tg128 3662.7 21.71 1118.3 tok/s 46.4 tok/s 6.435 656.4 tok/s 13.36 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 48.4 tok/s 1.00x 963.0 tok/s 963.0 tok/s 1063.3 3.722
2x 67.2 tok/s 1.39x 871.7 tok/s 435.9 tok/s 2349.2 6.159
4x 100.6 tok/s 2.08x 868.4 tok/s 217.1 tok/s 4569.1 9.804

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 70.0% 21 30 33 No
TRUTHFULQA 76.7% 23 30 15.3 No
GSM8K 96.7% 29 30 100.3 No
MATHQA 16.7% 5 30 69.6 No
HUMANEVAL 83.3% 25 30 172.8 No