can u do this 711 method to uncensored heretic qwen 3.6 35b moe version

#15
by allenhuggingface - opened

already test DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF , so far is the best qwen.. but for my 16gb vram still slow

For your VRAM maybe the maximum you could run is Q3. If your system is offloading to RAM your speeds will drop dramatically

@allenhuggingface - use the 9b variant David made, it beats the vanilla qwen3.6-35b. the 27B version fits with sub ~200k context on my 4090, (24gb vram) and kv at 8bit

@allenhuggingface

To answer the topic line:

Currently waiting for 4 bit compression to be able to train (on local hardware) the Qwen3.5/.6 35B-A3B models.
There are tickets/merges currently in progress in this regard.

ATM training must be done in 16 bit + the experts -> which require 80-100 GB of vram MIN.

@DavidAU
Hey, just wanted to say—I tried out your Qwen3.5-9B-MTP and I'm genuinely blown away. Running it on a 5060Ti with 16GB VRAM at Q4_K_M with a 256K context window, and honestly? It's already incredibly impressive.

Super excited to hear you're working on the 35B version (Qwen3.5 or 6)—huge fanboy moment here.

Quick question: any chance you could drop an NVFP4 quantized version of the 9B? It would definitely speed things up for folks like me.

And about the 35B MoE model—heard today that Qwen might be open-sourcing the Qwen3.8 series. Haven't verified it yet, but if there's a MoE version, that's a game-changer for running locally on consumer GPUs. Seriously, this would be amazing for the community.

Thank you so much.

RE: 3.7/3.8 ...
Watching daily for this , pins and needles here.

Sign up or log in to comment