Share your model speed here

#3
by anvme - opened

Hey everyone!

I think this topic could be really useful for users who have plenty of system RAM to run larger models, but are limited by GPU VRAM

To make comparisons easier and help others understand what models/settings might work for their hardware, please share your setup and performance using the format below:

  • 1. Model file / model name
  • 2. GPU: model + VRAM
  • 3. CPU: model
  • 4. RAM: size + type
  • 5. Generation speed: tokens/sec
  • 6. App / frontend: e.g. FreeToken, llama.cpp, LM Studio, etc.
  • 7. Launch parameters / settings:
  • 8. Extra notes: quantization, context size, offloading, optimizations, or anything else worth mentioning

This should make it much easier for people to compare setups and decide which quantization to download

Curious to see if my old M1 ultra will have enough juice to run GGUF.

Also surprised that GGUF are already being uploaded, I thought this was a new architecture. Does unsloth have a llama.cpp fork that can already run this?

Im hoping I can fit this in my AMD Ryzen Ai 395+ 128Gb sadly my 3 x 4090 will be useless for this LLM

It will fit on a 128GB UMA, just a matter quantization. Unsloth writes the 4-bit will be 110GB, so that might be a bit tight on 128GB, hopefully we can offload ngram to SSD (or stream them or whatever its called).

AI MAX+ 395 gfx1151 rocm7.14
Qwen3.8-Flash-Next-UD-IQ1_S
PP ~200-300t/s
TG ~20t/s

image

Not scientific benchmarks yet (just seeing the llama-server numbers as I test it)

M1 Ultra 128G
Qwen3.8-Flash-Next-UD-IQ1_S
PP ~ 400 tps
TG ~ 20 tps

Just a quick one before hitting bed:

AI MAX+ 395 gfx1151 + R9700 gfx1201
Qwen3.8-Flash-Next-UD-IQ4_XS

llama-bench --model /ai/models/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -p 512,4096 -n 128,1024 -d 0,512 --device Vulkan0/Vulkan1 -ngl 99 -fa 1 -ts 40/50 --split-mode layer

WARNING: radv is not a conformant Vulkan implementation, testing use only.
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon AI PRO R9700 (RADV GFX1201) (radv) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat
ggml_vulkan: 1 = AMD Radeon Graphics (RADV STRIX_HALO) (radv) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa dev ts test t/s
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 pp512 532.87 ± 11.64
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 pp4096 464.10 ± 4.44
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 tg128 23.25 ± 0.10
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 tg1024 23.03 ± 0.47
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 pp512 @ d512 497.23 ± 14.41
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 pp4096 @ d512 462.56 ± 7.05
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 tg128 @ d512 23.16 ± 0.46
qwen4exp A3B IQ4_XS - 4.25 bpw 87.24 GiB 176.94 B Vulkan 99 1 Vulkan0/Vulkan1 40.00/50.00 tg1024 @ d512 22.15 ± 0.22

build: 035e22731 (10656)

Sign up or log in to comment