Qwen 3.8 35B A3B Would be really good!

#161
by ShyliaSafetensors - opened

The overthink problem with dense 27B is kinda unbearable with low speed token generation. but A3B would be very fast and very efficent on consumer grade GPU's

great

I dont really think they will be doing that... almost as if they are doing something else?
(qwen4 soon?)

I've been thinking a lot about this. looking at the 27b CoT, I'm pretty sure most of it is kinda unfaithful, in the sense that it doesn't directly correspond to the reasoning it describes.
I recently gave qwen3.8-27b something with the acronym 'EY' and in the process of trying to recall who 'EY' is it mentioned elon musk, a few strange bits of grammar, and then finally arrived at eliezer yudkowsky. if this is the kind of compression and information density that's within 27b, I actually think the CoT will be much less legible on the 35b a3b version. it can be done, but it would be nice if we got some mech interp tooling to figure out what these models are actually doing (I'm thinking j/r-lenses, which would be really cheap for the qwen team to make)

incredible to have the equivalent of 2025 opus running locally on my hardware tho. I thought it would take a lot longer to reach this level

yeah ik it came fast...
but i think its going way faster than they can understand lol...

Yeah plz plz plz ๐Ÿ˜ญ๐Ÿ™ hope they do ๐Ÿฅบ

Honestly I might try to train a self improving llm... I am currently working on learning Abt that... soo...

Sign up or log in to comment