vs Instruct

#4
by Throghar - opened

Tested both, Thinking is performing far better as coding agent and delivering more complex results.
Also did several tricky and logical math questions to test abilities were done correctly on first try with thinking model.

Kinda making me doubt benchmark results vs instruct model which didnt work as well

Owner

Some uses cases work better with thinking VS instruct.
Frankly we were just as surprised here with INSTRUCT vs THINKING benchmarks too.

It maybe looping issues and/or length of thinking which is resulting in poor "thinking" benchmarks (IE too long thinking = fail).
However, fine tunes of "thinking" are improving the thinking benchmarks (the 7 ones we measure).

The difference between thinking/instruct benchmarks is consistent with different sized models.
And oddly ; larger model benchmarks are relatively close to smaller ones.
Another odd finding.

There are other issues with the Qwen3.5s also; which are being testing/tweaked - including base/root models.
These are under testing atm.

SIDE NOTE: we noticed this same issue with Qwen 3 Instruct VS Thinking.

Sign up or log in to comment