DavidAU's picture
Update README.md
54c8715 verified
|
Raw
History Blame Contribute Delete
4.14 kB
metadata
language:
  - en
  - zh
license: apache-2.0
tags:
  - MTP
  - fine tune
  - heretic
  - uncensored
  - abliterated
  - multi-stage tuned.
  - all use cases
  - thinking
  - reasoning
  - qwen3.6
  - coder
  - creative
  - writing
  - fiction
  - roleplaying
  - bfloat16
  - all use cases
pipeline_tag: image-text-to-text

40B : "There is a BIGGER storm coming... and it will take no prisoners."

1290 Tensors, 96 layers : 50% larger than 27B Qwen 3.6 it is based on.

ABOUT:

40B version(s) based on the wildly powerful "711" which operates in/near "OpenAI, Claude and Gemini" closed source intelligence levels:

https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

(Benchmarks of the "711" at the above repo that surpass the Qwen 3.6 27B in all respects)

TESTING:

  • Prototype 1, light tune for testing of TWO 700+ ARC-C (OpenAi,Claude, and Gemini level intelligence) Qwen 3.6B super tunes joined "at the hip" to make a 40B monster.
  • Prototype 4, Model "Fable-Fusion-711" joined "at the hip" with Deckard 27B (Qwen 3.6) to make a 40B monster in pre-tune testing.
  • Testing / benching and post expansion tuning in progress.

IMPORTANT:

  • THIS IS A WORK in PROGRESS, and will highlight some parts of this process (40B version of "711") as we proceed.
  • NAME of this repo will CHANGE as the project proceeds.
  • RUNNING Benchmarks [subject to change] below.

PROJECT NOTES:

  • If you want to join the waitlist, you will be notified by email when the FINAL, fully tested and optimized version releases.
  • During expansion is normal to lose some benchmarks levels, which are then restored during the post expansion tuning.
  • It will (likely) take a number of tuning rounds/steps/stages to bring the 40B up to "700" club status.

RUNNING NOTES [reverse order]:

  • Alpha 2 and 3, 3b in testing // additional "non trained" expanded also in testing.
  • Selecting model(s) for Beta staging in progress.
  • Prelim benchmarks for "Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha" (tuned expansion) posted, moving on to next stage(s). Other Alphas are pending too.
  • Benchmark below for "Qwen3.6-40B-Grand-Intelligence-One" (the root model for "Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha") BEFORE post expansion (27B to 40B) training.

BENCHMARKS by Nightmedia

           arc/c arc/e boolq hswag obkqa piqa  wino

Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha3b [mid-light repair, deeper tune only]
mxfp8      0.695,0.864,0.902,0.819,0.494,0.814,0.772
REMARKS: Slightly lower, however 100% stable. SOTA IQ/power at 40B.

Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha3 [mid-light repair tune only]
mxfp8      0.701,0.862,0.903,...
REMARKS: Excellent, but unstable with some prompts.

Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha2 [light repair tune only]
mxfp8      0.690,0.864,0.908,...

Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha [light repair tune only]
mxfp8      0.689,0.859,0.903,...

------------------------------------------------------------
RAW EXPANDED MODEL(s) - pre stage to be tuned/adjusted.
------------------------------------------------------------

Qwen3.6-40B-Grand-Intelligence-One 
[NOT TRAINED YET, expansion only]
mxfp8      0.675,0.860,0.900,0.795,0.478,0.801,0.752

Qwen3.6-40B-Grand-Intelligence-Four-raw
[NOT TRAINED YET, expansion only]
mxfp8      0.679,0.853,0.905

------------------------------------------------------------
ORG MODELS FROM QWEN, no tuning, non heretic.
------------------------------------------------------------

Qwen3.6-27B-Instruct: [base, non heretic]
mxfp8      0.647,0.803,0.910,0.773,0.450,0.806,0.742

Qwen3.6-35B-A3B-Instruct [base, non heretic]
mxfp8      0.581,0.757,0.892,0.751,0.428,0.803,0.688

Qwen3.5-27B-Instruct: [base, non heretic]
mxfp8      0.557,0.711,0.868,0.533,0.452,0.706,0.695

NOTES:

  • Models are tested in "Instruct" mode because this generally works better with the testing harness.
  • Testing via "thinking" mode also shows the metrics (and changes) but not the true extent.
  • In actual fact when the model IS in thinking mode, it will exceed INSTRUCT benchmark scores in most cases.