How to use from
Pi
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF:
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF:"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated

This is an uncensored version of Qwen/Qwen3.6-35B-A3B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. I made GGUF's in Q8 for almost perfect full quality and MXFP4 to fit on a single RTX 3090 or 4090.
You can also export GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 to enable unified memory on linux with Nvidia.
Both model quants tested with mmproj-BF16.gguf for multi-modal image use.

llama.cpp latest version

llama.cpp/build/bin/llama-server \
  -m Qwen3.6-35b-abliterated-Q8_0.gguf \
  --mmproj mmproj-BF16.gguf \
  --host 0.0.0.0 \
  --port 11434 \
  --api-key XXX \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0\
  --n-gpu-layers 99 \
  --split-mode layer \
  --no-mmap \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --repeat-penalty 1.0 \
  --repeat-last-n = 256 \
  --jinja \
  -c 262144 \
  -b 2048 \
  -ub 2048 \
  --parallel 1 

Usage Warnings

  • Risk of Sensitive or Controversial Outputs: This model’s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated outputs.

  • Not Suitable for All Audiences: Due to limited content filtering, the model’s outputs may be inappropriate for public settings, underage users, or applications requiring high security.

  • Legal and Ethical Responsibilities: Users must ensure their usage complies with local laws and ethical standards. Generated content may carry legal or ethical risks, and users are solely responsible for any consequences.

  • Research and Experimental Use: It is recommended to use this model for research, testing, or controlled environments, avoiding direct use in production or public-facing commercial applications.

  • Monitoring and Review Recommendations: Users are strongly advised to monitor model outputs in real-time and conduct manual reviews when necessary to prevent the dissemination of inappropriate content.

  • No Default Safety Guarantees: Unlike standard models, this model has not undergone rigorous safety optimization. huihui.ai nor I bear any responsibility for any consequences arising from its use.

Donation

Your donation helps us continue our further development and improvement, a cup of coffee can do it.
  • bitcoin:
  bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ah
Downloads last month
279
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF

Quantized
(772)
this model

Collection including andyjack/Huihui-Qwen3.6-35B-A3B-abliterated-GGUF