This is a MXFP4_MOE quantization of the model trohrbaugh / Ling-3.0-flash-heretic

Quick Start

  1. Download the latest release of llama.cpp.

Recommended parameters from inclusionAI:

  • temperature=0.6
  • top_p=0.95
  • top_k=20

Performance

Metric This model Original model
KL divergence 0.0526 0 (by definition)
Refusals 0/100 87/100
Downloads last month
594
GGUF
Model size
127B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for noctrex/Ling-3.0-flash-heretic-MXFP4_MOE-GGUF

Quantized
(3)
this model