Raullen commited on
Commit
fc59f62
·
verified ·
1 Parent(s): 05a2b87

bench: add M3 Ultra results via rapid-mlx 0.8.18

Browse files
Files changed (1) hide show
  1. README.md +17 -0
README.md CHANGED
@@ -42,3 +42,20 @@ print(generate(model, tokenizer, prompt="Hello", max_tokens=32))
42
 
43
  - This is a pure text-generation MLX release. No vision/image inputs.
44
  - For best chat behavior, use the chat template that ships with this repo.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42
 
43
  - This is a pure text-generation MLX release. No vision/image inputs.
44
  - For best chat behavior, use the chat template that ships with this repo.
45
+
46
+ ## Benchmarks
47
+
48
+ > Measured on M3 Ultra Studio (28 (20 Performance and 8 Efficiency) CPU, 60-core GPU, 256 GB unified memory) via rapid-mlx 0.8.18.
49
+
50
+ | Variant | Decode tok/s | TTFT (ms) | Prefill 1k (tok/s) | Prefill 4k (tok/s) | Prefill 16k (tok/s) | Tool-call e2e |
51
+ |---|---:|---:|---:|---:|---:|---:|
52
+ | Tmax-9B (8-bit MLX) | 67.5 | 143 | 1,054 | 1,121 | 1,086 | 871 ms (OK) |
53
+
54
+ Full results (all 7 Tmax MLX variants + 2 Qwen3.5 controls): [rapid-mlx docs](https://github.com/raullenchai/Rapid-MLX/blob/main/docs/benchmarks/tmax-m3-ultra.md).
55
+
56
+ Reproduce:
57
+
58
+ ```bash
59
+ pip install rapid-mlx==0.8.18
60
+ rapid-mlx serve tmax-9b-8bit --port 8765
61
+ ```