Qwen3-30B-A3B-Instruct-2507-REAM-heretic - GGUF

This repository contains GGUF format model files for sasa2000/Qwen3-30B-A3B-Instruct-2507-REAM-heretic.

These files were quantized using llama.cpp to make them runnable on consumer hardware via CPU and GPU offloading.

Model Architecture Details

  • Architecture: Qwen3 MoE (Mixture of Experts)
  • Total Parameters: ~30B
  • Active Parameters: ~3.3B (A3B)
  • Max Context Length: 256K (Note: Requires significant RAM/VRAM at full context; lowering to 8K-32K is recommended for standard hardware).

Available Quantizations

File Name Bit Depth Description
Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q4_K_M.gguf 4-bit โญ Recommended. Best balance of speed, RAM usage, and intelligence.
Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q5_K_M.gguf 5-bit Higher quality, slightly larger RAM requirement.
Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q6_K.gguf 6-bit Very high quality, near unquantized performance.
Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q8_0.gguf 8-bit Extremely high quality, large file size.
Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q3_K_M.gguf 3-bit Smallest file size, noticeable intelligence loss. Use only if heavily RAM constrained.

Prompt Format (ChatML)

This model uses the ChatML format. If you are using a UI like LM Studio or Ollama, it should detect this automatically from the GGUF metadata.

<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
Downloads last month
113
GGUF
Model size
23B params
Architecture
qwen3moe
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Abiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF