Loving this version

#4
by SaturnsVoid - opened

Running the Quality version at 128k on my 9070XT and 48GB ram and getting 45/ts its able to one shot a lot of GoLang coding prompts in Pi.dev, Just wish it was uncensored.

Owner

Running the Quality version at 128k on my 9070XT and 48GB ram and getting 45/ts its able to one shot a lot of GoLang coding prompts in Pi.dev, Just wish it was uncensored.

Thank you, I’m glad the Quality version is working well for you, especially at 128K context.

I’m currently working on an uncensored version of this model. It has been more difficult than expected, so I can’t promise the final result yet. My preference is still to complete the uncensoring myself, so I can fully control the processing and validation pipeline.

If I cannot reach an acceptable result, the fallback would be to start from a high-quality uncensored base released by another researcher, then add native MTP preservation and APEX quantization on top of it.

For now, I’m still trying to finish the uncensored version independently and verify that it does not introduce significant quality regressions.

Owner

Running the Quality version at 128k on my 9070XT and 48GB ram and getting 45/ts its able to one shot a lot of GoLang coding prompts in Pi.dev, Just wish it was uncensored.

Quick update: I actually finished the uncensored version much sooner than expected 😅

After a lot of testing and some custom Heretic modifications for the 35B Qwen3.5 MoE architecture, I was able to get a stable result, preserve native MTP, and complete the APEX GGUF release.

BF16 + MTP:
https://huggingface.co/SC117/Ornith-1.0-35B-Heretic-MTP

APEX GGUF:
https://huggingface.co/SC117/Ornith-1.0-35B-Heretic-MTP-APEX-GGUF

Since you mentioned using the Quality version at 128K, you may want to try the Heretic Quality build directly.

Dear Author,
I love the SC117/Ornith-1.0-35B-MTP-APEX-GGUF compact version (after test many model)
Does the Heretic version is better and faster,
Thanks you!

Dear Author,
I love the SC117/Ornith-1.0-35B-MTP-APEX-GGUF compact version (after test many model)
Does the Heretic version is better and faster,
Thanks you!

Thanks for your support!

The Heretic version is mainly focused on reducing refusals while keeping the original capabilities as much as possible. It is not necessarily faster than the normal version, because the speed mainly depends on the quantization profile and your hardware.

If you are already happy with the Compact version, the Heretic Compact build would be the closest comparison. The Quality/Balanced versions can provide better accuracy but require more memory.

Thanks again for testing!

Sign up or log in to comment