Eval of 711, Tess, AEON, ThinkingCap, 27B base with no confounds

#35
by tcclaviger - opened

Here ya go, objective, independent, evidence that on the whole 711 is hands down an upgrade:

https://blog.robai.net/27bevals/

The Q8_0 vs BF16 711 results were repeated multiple times because it's odd but the fact is Q8 sometimes snaps a result to a better outcome, just is what is.

I've finished initially implementing llama style DRY in vllm with MTP compatibility.

When I've had some time to make sure that is not causing issues, I'll retest all of them and see if it changes things.

@tcclaviger
I can not access link/report from by location ; geo (AU) and/or ip blocked?
Can you add screenshots in case others are having this issue too?
thank you ;

Ah right I forgot I have some....stuff setup that might cause that.

I'll export and post elsewhere.

DavidAU pinned discussion

I did loads of testing with claude created tests across various quants and models
https://claude.ai/code/artifact/4472b727-8ae6-4099-b1a0-700ecf74bf2e

Every model was K_M and I also threw in IQ4_XS for the final tests, run in llamacpp
defiant-9b
defiant-9b-mtp
fable-711-q4
fable-711-q4-mtp
fable-711-q5
fable-711-q5-mtp
qwythos-q4
qwythos-q4-mtp
qwythos-q5
qwythos-q5-mtp

711 was the best out of all them by far. It was weird though, it seemed like Q4 was better at some stuff than Q5 and vice versa.
I could upload the testing files if anyone really wanted them, but I doubt anyone does. it took like 5-6 days of solid testing to get these results using a 5090. nearly 24 hours for the "abyssal" tests, it was obnoxious trying to get sonnet to not score a 100% across the board

image

Sign up or log in to comment