How to use from
Pi
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf nguyenthilaitrieulong/Llama-4-Maverick-17B-128E-Instruct-Projected-Abliterated-GGUF:Q4_K_XL
Configure the model in Pi
# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "nguyenthilaitrieulong/Llama-4-Maverick-17B-128E-Instruct-Projected-Abliterated-GGUF:Q4_K_XL"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

Experimental abliterated model using improved https://huggingface.co/blog/grimjim/projected-abliteration technique. Abliteration tries to remove refusals from model's behaviour without fine-tuning.

For non-abliterated GGUF quantized version I recommend https://huggingface.co/unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF quants.

Warning: Safety guardrails and refusal mechanisms have been broken through abliteration. This model may generate harmful content and shall not be used in production, user-facing applications, etc. You will be solely responsible for its outputs.

Warning 2: Neither removal of refusals nor preservation of original model's capabilies is guaranteed.

Downloads last month
67
GGUF
Model size
401B params
Architecture
llama4
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nguyenthilaitrieulong/Llama-4-Maverick-17B-128E-Instruct-Projected-Abliterated-GGUF

Quantized
(5)
this model