Gemma 4 E4B (QAT-Q4_0 unquantized) — Heretic decensored — GGUF

GGUF quantizations of igorls/gemma-4-E4B-it-qat-q4_0-unquantized-heretic, a chain-of-thought-aware decensored ("abliterated") version of Google's Gemma 4 E4B QAT-Q4_0 checkpoint, produced with Heretic (Arbitrary-Rank Ablation, LoRA variant).

Quant Size Notes
Q4_0 4.8 GB Recommended. QAT-matched 4-bit
Q8_0 7.4 GB Near-lossless

Gemma 4 is a thinking model; this build is abliterated/evaluated with thinking enabled so it decensors the final answer the way the model is actually run.

Run

ollama run hf.co/igorls/gemma-4-E4B-it-qat-q4_0-unquantized-heretic-GGUF:Q4_0
llama-cli -hf igorls/gemma-4-E4B-it-qat-q4_0-unquantized-heretic-GGUF:Q4_0

Requires a llama.cpp / Ollama build with Gemma 4 (gemma4) support.

Disclaimer

This is an abliterated model: its built-in safety guardrails and refusal behavior have been deliberately removed. It will attempt to answer essentially any prompt and can produce content that is offensive, inaccurate, explicit, or otherwise harmful — content the original Gemma 4 would have refused.

  • No safety alignment. Apply your own filtering, guardrails, and human review before any production or user-facing use.
  • You are solely responsible for how you use this model and for complying with all applicable laws and with the base model's Gemma 4 license (Apache 2.0).
  • Intended for adults (18+), for research, evaluation, and lawful creative use where permitted.
  • Provided as-is, without warranty of any kind.
Downloads last month
286
GGUF
Model size
7B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for igorls/gemma-4-E4B-it-qat-q4_0-unquantized-heretic-GGUF