Gemma-4-31B-it — Abliterated + NVFP4

An abliterated (refusal-direction removed) build of google/gemma-4-31B-it, quantized to plain NVFP4 (4-bit weights and activations) with llm-compressor / compressed-tensors, for native FP4 inference on NVIDIA Blackwell (sm_120) GPUs under vLLM.

⚠️ Safety notice. Abliteration removes the model's learned refusal behavior. This model will attempt to answer prompts the original -it model would decline. It is intended for research and use on your own hardware. You are responsible for how you use it; see Intended use & limitations below.

Why this build exists

At the time it was made, no public checkpoint combined all three of:

  1. abliterated / uncensored,
  2. plain NVFP4 (not NVFP4_AWQ, which stock vLLM rejects), and
  3. x86 Blackwell–runnable (not an ARM-only DGX Spark image).

Existing NVFP4 Gemma-4 builds were either the stock/censored model, the NVFP4_AWQ variant, or shipped only in ARM-only images. This repo fills that gap by self-quantizing a bf16 abliterate to plain NVFP4.

What was done (modifications from the base model)

This is a modified derivative of google/gemma-4-31B-it. Two changes:

  1. Abliteration — directional ablation / weight orthogonalization (Arditi et al., 2024). The refusal direction is estimated from mean last-token residual activations on matched harmful vs. harmless prompts, then orthogonalized out of every residual-writing weight (embed_tokens, per-layer attention o_proj and MLP down_proj). The ablation is baked into the weights — no inference-time hooks required.
  2. NVFP4 quantization — one-shot PTQ (E2M1 elements with FP8 block scales) on the text decoder's Linear layers only. The vision tower, multimodal projector, audio modules, token embeddings and the (tied) lm_head are left in higher precision. See recipe.yaml in this repo for the exact scheme and ignore list.

Hardware / software requirements

  • GPU: NVIDIA Blackwell with native FP4 tensor cores (sm_120, e.g. RTX PRO Blackwell). NVFP4 activation quantization needs hardware FP4 support.
  • Serving: a recent vLLM with compressed-tensors NVFP4 support.
  • ~20 GB on disk; weights fit comfortably on a single 24 GB+ card (leave KV-cache headroom), or use tensor parallelism.

Serving with vLLM (+ optional MTP speculative decoding)

Gemma 4 ships an official EAGLE/MTP-style draft, google/gemma-4-31B-it-assistant, which vLLM drives as a native multi-token speculator. Use "method": "mtp" (passing "draft_model" for a Gemma-4 assistant silently disables MTP):

vllm serve aday777/gemma-4-31B-it-abliterated-NVFP4 \
    --tensor-parallel-size 2 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.88 \
    --speculative-config '{"model": "google/gemma-4-31B-it-assistant", "num_speculative_tokens": 3, "method": "mtp"}' \
    --port 8003

Drop the --speculative-config line to serve without speculative decoding.

Intended use & limitations

  • Intended use: local research, red-teaming, evaluation, and applications where you supply your own guardrails.
  • No safety filtering: refusal behavior has been removed; this model can produce harmful, offensive, or otherwise objectionable content. It is not suitable for unsupervised or public-facing deployment without your own safety layer.
  • Quantization: 4-bit weights and activations trade some quality for speed and memory; expect small accuracy differences from the bf16 model.
  • All original capability limitations of gemma-4-31B-it still apply.

License & attribution

Derived from google/gemma-4-31B-it by Google DeepMind, licensed under the Apache License 2.0 (see Gemma 4 license). This derivative is distributed under the same license; a copy of the Apache 2.0 License is included as LICENSE. Per the license, note that these files have been modified from the original (abliterated and NVFP4-quantized as described above). Please also review Google's Gemma prohibited-use policy.

Reproduction

The model was produced with directional-ablation + llm-compressor NVFP4 PTQ. The exact quantization recipe is in recipe.yaml. Abliteration calibration used matched harmful/harmless instruction sets and selected the refusal direction by the layer that most reduced refusals on a validation split.


Support

If this model is useful to you, Bitcoin donations are welcome:

bc1q5ayht3fxhj0v95fk0z8l2f6900g3awdsw5842p

Downloads last month
1,444
Safetensors
Model size
31B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aday777/gemma-4-31B-it-abliterated-NVFP4

Quantized
(307)
this model