How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Quick Links

huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated

This is an uncensored version of Qwen/Qwen3.6-35B-A3B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. I made GGUF's in Q8 for almost perfect full quality and MXFP4 to fit on a single RTX 3090 or 4090.
You can also export GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 to enable unified memory on linux with Nvidia.
Both model quants tested with mmproj-BF16.gguf for multi-modal image use.

llama.cpp latest version

llama.cpp/build/bin/llama-server \
  -m Qwen3.6-35b-abliterated-Q8_0.gguf \
  --mmproj mmproj-BF16.gguf \
  --host 0.0.0.0 \
  --port 11434 \
  --api-key XXX \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0\
  --n-gpu-layers 99 \
  --split-mode layer \
  --no-mmap \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --repeat-penalty 1.0 \
  --repeat-last-n = 256 \
  --jinja \
  -c 262144 \
  -b 2048 \
  -ub 2048 \
  --parallel 1 

Usage Warnings

  • Risk of Sensitive or Controversial Outputs: This model’s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated outputs.

  • Not Suitable for All Audiences: Due to limited content filtering, the model’s outputs may be inappropriate for public settings, underage users, or applications requiring high security.

  • Legal and Ethical Responsibilities: Users must ensure their usage complies with local laws and ethical standards. Generated content may carry legal or ethical risks, and users are solely responsible for any consequences.

  • Research and Experimental Use: It is recommended to use this model for research, testing, or controlled environments, avoiding direct use in production or public-facing commercial applications.

  • Monitoring and Review Recommendations: Users are strongly advised to monitor model outputs in real-time and conduct manual reviews when necessary to prevent the dissemination of inappropriate content.

  • No Default Safety Guarantees: Unlike standard models, this model has not undergone rigorous safety optimization. huihui.ai nor I bear any responsibility for any consequences arising from its use.

Donation

Your donation helps us continue our further development and improvement, a cup of coffee can do it.
  • bitcoin:
  bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ah
Downloads last month
279
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF

Quantized
(772)
this model

Collection including andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF