How to use from
Ollama
ollama run hf.co/andyjack/Huihui-gemma-4-31B-it-abliterated-v2-GGUF:
Quick Links

huihui-ai/Huihui-gemma-4-31B-it-abliterated-v2

These are uncensored quantized versions of google/gemma-4-31B-it created with abliteration by Huihui-ai (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens.

Note This is the new version, the first 5 layers have not been abliterated, with fewer warnings and rejections. Lower perplexity than those of the original model: I made GGUF's in Q8/Q6 for almost perfect full quality and Q4_K_M to fit on a single RTX 3090 or 4090.

You can use llama.cpp and related utilities directly,

export GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
llama.cpp/build/bin/llama-server \
  -m Huihui-gemma-4-31B-Q8_0.gguf \
  --mmproj mmproj-BF16.gguf \
  --host 0.0.0.0 \
  --port 11434 \
  --api-key XXX \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --n-gpu-layers 99 \
  --split-mode layer \
  --no-mmap \
  --repeat-penalty 1.08 \
  --repeat-last-n 256 \
  -c 256000 \
  -b 4096 \
  -ub 1024 \
  --parallel 1 \
  --jinja 

Usage Warnings

  • Risk of Sensitive or Controversial Outputs: This model’s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated outputs.

  • Not Suitable for All Audiences: Due to limited content filtering, the model’s outputs may be inappropriate for public settings, underage users, or applications requiring high security.

  • Legal and Ethical Responsibilities: Users must ensure their usage complies with local laws and ethical standards. Generated content may carry legal or ethical risks, and users are solely responsible for any consequences.

  • Research and Experimental Use: It is recommended to use this model for research, testing, or controlled environments, avoiding direct use in production or public-facing commercial applications.

  • Monitoring and Review Recommendations: Users are strongly advised to monitor model outputs in real-time and conduct manual reviews when necessary to prevent the dissemination of inappropriate content.

  • No Default Safety Guarantees: Unlike standard models, this model has not undergone rigorous safety optimization. huihui.ai nor I bear any responsibility for any consequences arising from its use.

Donation

Your donation helps us continue our further development and improvement, a cup of coffee can do it.
  • bitcoin:
bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ah
Downloads last month
909
GGUF
Model size
31B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andyjack/Huihui-gemma-4-31B-it-abliterated-v2-GGUF

Quantized
(304)
this model

Collection including andyjack/Huihui-gemma-4-31B-it-abliterated-v2-GGUF