ArtificialAnalysis score of 52 outscore GLM 5.2 , Opus 4.6 and touches Opus 4.7!!!!

#143
by mayankiit04 - opened

This is unbelievable!!

It doesn't come close to approaching the overall performance of GLM 5.2, Opus 4.6 etc. AA is fundamentally flawed (averages redundant tests, multiple choice etc.), and Alibaba test maxes (much better on the tests than real world scenarios). Keep an a eye on Arena. I'm confident Qwen 3.8 27b will only be around Qwen 3.5 27b, and much lower than Gemma 4 31b, across most domains.

Well what prevents the larger LLMs from not benchmaxxing when $Trillions are at stake... I find this argument to be very vague ...

@mayankiit04 It's not as simple as there's trillions at stake, so let's all benchmax. If the big Western companies like Google, OpenAI and Anthropic did the same the damage to their reputations would be very expensive.

For example, the English SimpleQA is a non-multiple choice broad English knowledge test, consequently, the scores reliably correlate with parameter count. There's simply no theoretical way to legitimately boost the scores with fine-tuning or any other post training technique. The only way is to up the parameter count and/or fully train on a broad English language corpus.

Anyways, the 70b- leader is Llama 3 70b with a score of ~20 (same goes for the updated English SimpleQA version), and Gemma 4 31b scores ~11. Yet Qwen 3.5 27b, which has the broad English knowledge of models that scored <10, scored an impossible 25, which only 200b+ models have ever legitimately achieved.

If Google tried to claim a 25 English SimpleQA score with Gemma 4 31b, or Mistral with their Mistral Small series, and so on, it would shock the industry. But with Chinese companies it's par for the course.

okay... i get some technical point.. but i still feel from an economical point of view, every big llm can be benchmaxxing as otherwise will hurt their pr.. i have not heard much about arena but AA seems quite solid to me.

@mayankiit04 It's not that AA isn't solid, and it's definitely aboveboard. But it relies on static multiple choice tests, which is highly problematic, even without deliberate cheating, since subpar contamination mitigation can notably inflate scores. And thinking makes things worse by increasing the chance that the contamination will make it into the context and artificially boost the score.

If you check out the arena (https://arena.ai/leaderboard/text) you can see that across all categories Gemma 4 31b is ranked notably higher than Qwen3.5 27b (rank of 45 vs 116).

Take creative writing for example. Qwen3.5's stories are more repetitive, are filled with more contradictions to itself and the user's prompt, experience a much larger drop in quality when forced to write original stories that respect a list of inclusions and exclusions, and so on. Another example is thinking. Qwen 3.5+ burns through a lot more tokens, goes off on more irrelevant tangents, is more likely to get stuck in infinite loops, and so on. Plus in general the instruction following is far worse, with Qwen 3.5 trying to return a nearest match vs respond to the nuance in the user's prompt. It's not my personal opinion. Despite it's higher tests scores, including AA, Qwen 3.5 is a MUCH weaker AI model. This is even more true with their smaller models (e.g. 4b). They're absurdly weak, yet have test scores of models ~10x their size. Alibaba strongly favors testmaxing over real-world performance in all but a handful of domains (e.g. coding).

Interesting. But honestly, that claim about Qwen3.8 27B being neck and neck with Opus 4.6 Max? Yeah, we still need to see some receipts from independent benchmarks. From my own testing, Gemma 4 31B actually feels way more cracked, especially when it comes to coding, compared to Qwen3.8 27B. Qwen definitely takes the W in tool calling, but once you fine-tune Gemma, its tool calling goes absolutely hard too.

Something doesn't add up. Qwen 3.8 27b has a much higher Artificial Analysis score than Gemma 4 31b (52 vs 30), but on the Arena it performs notably worse across domains than Gemma 4 31b (final ranking of 81 vs Gemma 4 31b's 62), and far worse than the models with comparable ~50 AA scores.

There are several major issues with most tests, including AA. For example, the questions can't be repeated. This is the only way to rule out cheating and contamination. A new batch of equally challenging questions will score approximately the same, which you can confirm be re-testing some previously scored models. The questions must also be one-shot and worded various ways, which include being riddled with errors (e.g. grammatical and spelling) and irrelevant data so the LLM is forced to see past the noise, just like in real-life use cases. The answers must also be fully retrieved (not multiple choice). Users rarely give a list of possible correct answers for the LLM to choose from, plus this rewards overtraining since weakly held data can be enough to get a few more multiple choice questions right, but at the same time get more real-world fully retrieved answers wrong, spiking hallucinations in real-world use cases.

Anyways, one thing I'm absolutely certain of is Qwen3.6 27b ain't no Opus, GPT or Gemini. It doesn't come close. So if AA thinks they're comparable then either AA is a complete failure as a metric or Alibaba cheated/testmaxed.

@phil111 spot on fr fr. honestly i feel like automated benchmarks like AA are becoming kinda useless nowadays. companies are just test maxing and teaching their models to guess the answers instead of actually making them smarter. the arena is basically a vibe check from real humans, and people care way more about formatting, tone, and actual helpfulness, which gemma 4 clearly nails.

those multiple choice benchmarks just reward overtrained models that completely fumble in real world use lol. we seriously need better benchmarks for real world use cases, and obviously it has to be neutral.

@Ashacorporation That's certainly true about the arena. I've seen examples where the user upvoted the unambiguously wrong answer because they favored the style and politeness. Ideally users should only prompt about things they're experts in (engineers about engineering, medical doctors about medicine, and so on). It baffles me that users are testing models on the arena about things they're clueless about.

This is unbelievable!!

Why don't you test it on your own set of tasks instead of blindly relying on benchmarks? I can say that the GLM-5.2@iq3_s I use significantly outperforms the Qwen3.8-27B on my tasks.

Sign up or log in to comment