any updates on its 27b qwen 3.6 cousin

#51
by SolsticeAI - opened

hey so i came here from the discussion on the original model i just wanted to ask about the qwen 3.6 27b finetune variant ur making can we get some updates and details on that thx so much for ur prev. response btw

yeah, i'd rather have a 35b moe model but i was also wondering what the progress on the models is

hey yes ill second that lol

I dream of someone taking Qwen 3.6 27B and pairing it down to a Qwen 3.6 9b-12b coder so I can use them on smaller systems when I don't want to use my main workstations or cloud compute.

bro you stole the words from my mouth i too dream of such paradise

I personally look for coding models. It seems there might be a time when top LLMs would be too expensive or unavailable (regionally or worldwide) and we'll be desperate for local ones. If this kind hybrid model wouldn't have wikipedia inside I don't mind πŸ˜‰ I needs to be fairly fast, good reasoning and know how to work with tools. My employer doesn't like the idea of dropping a few grands on hardware just to play around. And I'm stuck to my GTX 1660... Yes, A3B/A4B models produce 5 t/s max and I leave 30Bs to work overnight πŸ˜† So I'm waiting for a stable version of this model!

I can't imagine using llm's for work and my employer either not paying for reasonable hardware or paying for api access. I'd probably stick to free api's online or just tell the employerer they need to up their budget or cut ai for now. (Model size I find doesn't matter if you can't use a reasonable context size, especially for agentic code)

i mean turboquant fits a 1 mil or 262k context into like 2-4 gigs. extra MAX losslessly so context meh

Owner

Hey @shreyan35 @TESTPOINTrxz @bleaki @ka4ep β€” thanks for your patience, and for following over from the main model. Quick status:

Gemma 4 12B v3 is basically at the door. It's been through 4–5 internal iterations β€” this one fixes the known issues people have been reporting and adds some capability on top. I'm expecting to ship it in the next day or two.

Qwen 3.6 27B is targeted for next weekend. Honestly, the reason it's not already out is that the 12B fix turned out much harder than I expected and ate most of my time. But I've already got a large batch of new Fable data prepped and ready for the 27B, so I don't expect big trouble there once I switch over.

On the smaller-hardware / distilled-coder wishes (@bleaki @ka4ep ) β€” I hear you, and that "fast, tool-capable, doesn't need to know all of Wikipedia" local coder is exactly the kind of thing I care about too. I can't put a 9–12B pared-down version or a 35B MoE finetune on a firm timeline yet (paring down and training MoE on a single-GPU setup is genuinely hard), but it's on my radar and I'll keep the low-VRAM crowd in mind. Updates will land right here β€” stay tuned.

If it's any use / there's a way to remotely use it, i've got an extra 4070 that sits idle in a pc 95% of the time.
I know it's only 12gb vram, but that's why it's sitting idle - if you can take advantage of it though let me know @yuxinlu1

thank you so much @yuxinlu1 will the performance be equivalent to qwen 3.6 27b please do let me know, ive got some very exciting ideas with this model

Sign up or log in to comment