Great model so far

#17
by mythrime - opened

Tested so far, it kinda solve the overthinking, althought sometimes it still think for 30 minutes or so ( probably because I set to xhigh ), but the quality of 30 minutes thinking under this model surpass the stock model under the same xhigh thinking duration, the solution it produce even beats out cloud model im subscribing for and i was thinking to unsubs them and use this as my primary model as well, great job!

Waiting for the Fable Fusion version, thanks again!

Thank you so much.
If I may ask ; what prompt[s] provoked 30 mins of thinking?

This will help in the test lab / refining future models.
We have a list tests we use to "stress test" the models.

If im going to be direct like code explanation, or code validation, it mostly done around 5-7 minutes.

the 30 minutes mostly happen when i asks to read the code for optimisation, like asking in a formatted way such as "Given {featureName} feature, and the code is located on file A,B,C, in general it does {featureBLL}, but found issue are {issues[]}, try fix the bug and optimise it, expected optimisation are {successCondition}"

the thinking got longer, usually up to 1h the longer the {successCondition} are so i believe it just try to find some extra possibility, which so far giving me code quality that exceed my expectation so far. but yea those last line probably the one causing it

The LOW version of IQ4_XS fits on my 7900XT with 128k context. Nice!

DavidAU pinned discussion

The LOW version of IQ4_XS fits on my 7900XT with 128k context. Nice!

If you’re looking for a version with limited video memory, you might be interested in this model:
https://huggingface.co/Bucoid/Qwen3.8-27B-Heretic-Ara-16GB-VRAM-IQ4-XS-MTP-GGUF
The Ara method (actively promoted by the author Bucoid) is a highly optimized custom quantization pipeline that extreme compression and inference acceleration for a specific amount of video memory (16 GB VRAM).

Thank you so much.
If I may ask ; what prompt[s] provoked 30 mins of thinking?

This will help in the test lab / refining future models.
We have a list tests we use to "stress test" the models.

@DavidAU - i know i asked before, but can you please make a finetune that retains the same level of xhigh reasoning/thinking as the base/unsloth version? please review the AA analysis of what proper xhigh does when it is not hampered.

https://artificialanalysis.ai/models/open-source/small <- score of 52 on intelligence.
https://artificialanalysis.ai/models/open-source/large <- if you compare that to frontier models, it matches deepseek v4 0731 and is one point lower than glm-5.2 both at max reasoning.

For those of us who want/need the xhigh for complicated codebases along with your incredible finetuning, i beg and implore you to leave the xhigh reasoning as is and even enhance it. i'd happily wait extra time to get through more complicated problems if i swap to xhigh reasoning or run xhigh as my daily driver, while others with less time or gpu resources and stick with low or medium reasoning.

Sign up or log in to comment