Agents-A1-oQ2

This model was quantized using oQ (oMLX v0.5.3) mixed-precision quantization.

Base model: Agents-A1

Quantization details

  • Model type: qwen3_5_moe
  • Bits: 2
  • Group size: 64
  • Format: MLX safetensors

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.5.3
  • Max Concurrent Requests: 4
  • Settings:
    • Thinking: Disabled
    • TurboQuant KV Cache: Enabled

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 1059.4 20.91 966.6 tok/s 48.2 tok/s 3.728 309.0 tok/s 12.59 GB
pp4096/tg128 3665.4 21.35 1117.5 tok/s 47.2 tok/s 6.400 660.0 tok/s 13.30 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 48.2 tok/s 1.00x 966.6 tok/s 966.6 tok/s 1059.4 3.728
2x 67.0 tok/s 1.39x 864.7 tok/s 432.4 tok/s 2368.4 6.192
4x 97.5 tok/s 2.02x 856.3 tok/s 214.1 tok/s 4625.5 10.032

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 83.3% 25 30 29.4 No
TRUTHFULQA 83.3% 25 30 15.1 No
GSM8K 90.0% 27 30 93.3 No
MATHQA 46.7% 14 30 17.5 No
HUMANEVAL 83.3% 25 30 158 No
Downloads last month
219
Safetensors
Model size
4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-works/Agents-A1-oQ2