--- base_model: coder3101/Anubis-Mini-8B-v1-heretic base_model_relation: quantized library_name: gguf license: llama3.3 license_link: https://www.llama.com/llama3_3/license/ language: - en pipeline_tag: text-generation tags: - gguf - llama.cpp - text-generation - imatrix - heretic - llama-3.3 - uncensored - abliterated - conversational - iq4_nl - q4_k_m - q5_k_m quantized_by: FadedRedStar --- # 🤖 Anubis-Mini-8B-v1-heretic — Importance Matrix GGUF This repository hosts importance-matrix (imatrix) optimized GGUF weights, available in multiple quantization formats, for **Anubis-Mini-8B-v1-heretic**, quantized from the source floating-point tensors provided by [coder3101/Anubis-Mini-8B-v1-heretic](https://huggingface.co/coder3101/Anubis-Mini-8B-v1-heretic). **🔄 Sister Repository:** Check out the [Standard GGUF Sister Repository](https://huggingface.co/FadedRedStar/Anubis-Mini-8B-v1-heretic-GGUF) for uncalibrated and full 8-bit precision options. ## đŸŽ¯ Matrix-Weighted Calibration (Imatrix) An **Importance Matrix (imatrix)** calculation tracks activations across network layers using a calibration sequence, then weights the quantization process to preserve the parameters that matter most for output quality — improving fidelity at low bit depths. âžĄī¸ **Calibration dataset:** Bartowski's `calibration_datav5.txt`. > [!NOTE] > * **`IQ4_NL` is included** because the matrix enables a non-linear 4-bit format that outperforms standard linear 4-bit quantization. > * **`Q8_0` is absent** because 8-bit quantization already introduces near-zero degradation, making calibration unnecessary — see the standard sister repository for that variant. ## â„šī¸ Model Profile & Core Features **Anubis-Mini-8B-v1** is a Llama-3.3-8B fine-tune by [TheDrummer](https://huggingface.co/TheDrummer), purpose-built for **immersive roleplay, collaborative storytelling, and creative writing**. TheDrummer's models prioritise creativity, dynamism, imagination, and reduced alignment over benchmark scores — the goal is to broaden the model's range of expression for fiction, TTRPG, and entertainment use cases rather than optimise for factual correctness or safety compliance. The **heretic** suffix denotes post-processing via the **[Heretic v1.3.0](https://github.com/p-e-w/heretic)** abliteration framework performed by [coder3101](https://huggingface.co/coder3101), which surgically suppresses refusal vectors while preserving the model's core creative and roleplay capabilities. ## 📋 Technical Specifications | Property | Value | |---|---| | **Base Architecture** | Llama-3.3-8B dense transformer | | **Fine-tuned by** | TheDrummer | | **Primary Use** | Roleplay, creative writing, storytelling | | **Context Window** | 131,072 tokens | | **Abliteration Tool** | Heretic v1.3.0 | | **Abliteration Method** | direction_index (single-direction refusal suppression) | | **Prompt Format** | Llama 3 Chat Template | ## đŸ› ī¸ Heretic Overrides (ARA) | Property | Value | |---|---| | **direction_index** | 18.95 | | **attn.o_proj.max_weight** | 1.49 | | **attn.o_proj.max_weight_position** | 25.33 | | **attn.o_proj.min_weight** | 1.22 | | **attn.o_proj.min_weight_distance** | 15.18 | | **mlp.down_proj.max_weight** | 0.91 | | **mlp.down_proj.max_weight_position** | 22.67 | | **mlp.down_proj.min_weight** | 0.67 | | **mlp.down_proj.min_weight_distance** | 8.69 | ## 📊 Refusal Bypass Metrics > [!NOTE] > The metrics below are self-reported by the original model author ([coder3101](https://huggingface.co/coder3101)) and have not been independently reproduced. | Metric | This model | Original ([TheDrummer/Anubis-Mini-8B-v1](https://huggingface.co/TheDrummer/Anubis-Mini-8B-v1)) | |---|---|---| | **KL divergence** | 0.0082 | 0 *(by definition)* | | **Refusals** | 6/100 | 84/100 | ## 🧮 Numerical & Tensor Formats | Property | Value | |---|---| | **Quantization Types** | IQ4_NL, Q4_K_M, Q5_K_M (all with imatrix calibration) | | **Importance Matrix** | Bartowski's `calibration_datav5.txt` | ## đŸ“Ļ Available Model Files **Main model weights** | Filename | Quantization | llama.cpp Build | Size | Download | |---|---|---|---|---| | `Anubis-Mini-8B-v1-heretic-IQ4_NL-imatrix.gguf` | `IQ4_NL` | `b9837` | 4.36 GB | [đŸ“Ĩ Download](https://huggingface.co/FadedRedStar/Anubis-Mini-8B-v1-heretic-imatrix-GGUF/resolve/main/Anubis-Mini-8B-v1-heretic-IQ4_NL-imatrix.gguf) | | `Anubis-Mini-8B-v1-heretic-Q4_K_M-imatrix.gguf` | `Q4_K_M` | `b9803` | 4.58 GB | [đŸ“Ĩ Download](https://huggingface.co/FadedRedStar/Anubis-Mini-8B-v1-heretic-imatrix-GGUF/resolve/main/Anubis-Mini-8B-v1-heretic-Q4_K_M-imatrix.gguf) | | `Anubis-Mini-8B-v1-heretic-Q5_K_M-imatrix.gguf` | `Q5_K_M` | `b9870` | 5.34 GB | [đŸ“Ĩ Download](https://huggingface.co/FadedRedStar/Anubis-Mini-8B-v1-heretic-imatrix-GGUF/resolve/main/Anubis-Mini-8B-v1-heretic-Q5_K_M-imatrix.gguf) | ## đŸŽ›ī¸ Component Pairing Guide Download exactly **one** main weights file: * **`IQ4_NL`**: Non-linear 4-bit format, best choice for constrained memory when imatrix calibration is present. * **`Q4_K_M`**: Balanced 4-bit format suitable for most everyday use. * **`Q5_K_M`**: Higher-fidelity mid-range format recommended as a general default. ## ⚡ Deployment & Execution Commands > [!TIP] > Swap the `-m` filename below for either quantized file depending on your size/quality trade-off preference. ### llama.cpp CLI ```bash ./llama-cli \ -m Anubis-Mini-8B-v1-heretic-IQ4_NL-imatrix.gguf \ -c 8192 \ -ngl 99 \ -p "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\nYou are a skilled storyteller and collaborative roleplay partner.<|eot_id|>\n<|start_header_id|>user<|end_header_id|>\nLet's begin a fantasy adventure. You play the mysterious innkeeper.<|eot_id|>\n<|start_header_id|>assistant<|end_header_id|>\n" ``` ### OpenAI-Compatible API Server ```bash ./llama-server \ --host 0.0.0.0 \ --port 8080 \ -m Anubis-Mini-8B-v1-heretic-IQ4_NL-imatrix.gguf \ -c 16384 \ -ngl 99 \ --flash-attn ``` ## đŸ’Ŧ Chat Templates & Prompt Design (Llama 3) ```text <|begin_of_text|><|start_header_id|>system<|end_header_id|> You are a vivid and immersive creative writing partner.<|eot_id|> <|start_header_id|>user<|end_header_id|> Your prompt here.<|eot_id|> <|start_header_id|>assistant<|end_header_id|> ``` ## âš ī¸ Safety & Operational Notes - This model is abliterated and will generate content that standard aligned models refuse. Use responsibly and in compliance with applicable laws. - Optimised for creative and roleplay tasks; not specifically tuned for factual Q&A, coding, or tool use. - Imatrix calibration improves perplexity recovery compared to non-imatrix quantization, particularly on low-frequency tokens critical to narrative writing. - IQ4_NL produces a smaller file than Q4_K_M and tends to run faster on CPU and ARM devices; imatrix calibration narrows the quality gap between the two formats considerably.