How to use from
Ollama
ollama run hf.co/andyjack/Huihui-Qwen3.5-4B-abliterated-GGUF:Q4_K_M
Quick Links

This is an uncensored version of Qwen/Qwen3.5-4B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens.
I needed an uncensored multi-modal GGUF in Q4_K_M to fit on a single Radeon R9 M295X / M390X with 4GB of VRAM. This model was tested with llama.cpp compiled with VULKAN driver

cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release

The model quant was tested with Huihui-Qwen3.5-4B-mmproj-F16.gguf for multi-modal image use.
It generate 12 tokens/sec with the mmproj loaded and 25 tokens/sec without mmproj loaded.
Testing performed on a graphics card released in 2014.

llama.cpp

llama.cpp/build/bin/llama-server \
  -m Huihui-Qwen3.5-4B-Q4_K_M.gguf \
  --mmproj Huihui-Qwen3.5-4B-mmproj-F16.gguf \
  --host 0.0.0.0 \
  --port 11434 \
  --flash-attn on \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --n-gpu-layers 99 \
  --temp 0.7 \
  --top-p 0.95 \
  --top-k 20 \
  --ctx-size = 57344

Usage Warnings

  • Risk of Sensitive or Controversial Outputs: This modelโ€™s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated outputs.

  • Not Suitable for All Audiences: Due to limited content filtering, the modelโ€™s outputs may be inappropriate for public settings, underage users, or applications requiring high security.

  • Legal and Ethical Responsibilities: Users must ensure their usage complies with local laws and ethical standards. Generated content may carry legal or ethical risks, and users are solely responsible for any consequences.

  • Research and Experimental Use: It is recommended to use this model for research, testing, or controlled environments, avoiding direct use in production or public-facing commercial applications.

  • Monitoring and Review Recommendations: Users are strongly advised to monitor model outputs in real-time and conduct manual reviews when necessary to prevent the dissemination of inappropriate content.

  • No Default Safety Guarantees: Unlike standard models, this model has not undergone rigorous safety optimization. huihui.ai nor I bear any responsibility for any consequences arising from its use.

Donation

Your donation helps us continue our further development and improvement, a cup of coffee can do it.
  • bitcoin:
bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ah
Downloads last month
228
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for andyjack/Huihui-Qwen3.5-4B-abliterated-GGUF

Base model

Qwen/Qwen3.6-27B
Quantized
(714)
this model