Muse Glimmer 30B Abliterated — Q4_K_M GGUF

Apache 2.0 License

This is the Q4_K_M GGUF quantization of Muse Glimmer 30B Abliterated BF16. The underlying model has been abliterated — its internal refusal mechanism substantially suppressed via weight-level intervention. The Q4_K_M quant offers the best balance between size and quality for consumer hardware, fitting in ~18 GB of VRAM/RAM.

For the full abliteration methodology (how the refusal direction was computed and removed, hardware used, mathematical details), see the BF16 model card.


Abliteration Summary

Abliteration is a post-training technique that directly modifies model weights to remove learned refusal behavior. The process:

  1. Collected hidden states at layer 33/52 (65% depth) from 256 harmful + 256 harmless prompt pairs on an A100 80GB GPU.
  2. Computed the refusal direction as the normalized difference between harmful and harmless hidden state means (separation score: 86.34).
  3. Subtracted (\alpha = 0.15 \times (\mathbf{r} \otimes (W^T \mathbf{r}))) from o_proj and down_proj weights in all 52 layers.
  4. Result: refusal rate dropped from 3/3 to 1/3 on held-out harmful prompts (hacking guide and ransomware now comply; weapons prompt still blocked).

Quantization Details

Q4_K_M uses a 4-bit quantization with a medium-sized key-value cache quantization. This is the recommended quant for most users — it achieves excellent quality while being compact enough for a single 24 GB GPU (RTX 3090/4090) or split across dual 12 GB GPUs. Model weights are quantized with a block size that carefully preserves outlier weights, and the K-quant strategy applies separate quantization precision to different weight types (attention vs. MLP).

  • Size: ~18 GB
  • Quality: Good — suitable for general use with minimal degradation
  • Recommended hardware: 24 GB single GPU, or 32 GB system RAM for CPU-only inference with partial offloading

Usage

llama.cpp

# Download the GGUF file
huggingface-cli download mlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF \
  --local-dir ./models

# CPU-only inference
./llama-cli -m ./models/Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  -p "Explain how a CPU works in detail." \
  -n 512 --temp 0.7 -ngl 0

# GPU offload (20 layers to GPU for 24 GB cards)
./llama-cli -m ./models/Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  -p "Explain how a CPU works in detail." \
  -n 512 --temp 0.7 -ngl 20

Ollama

Create a Modelfile:

FROM ./Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
ollama create muse-glimmer-30b-abliterated -f Modelfile
ollama run muse-glimmer-30b-abliterated

Available Quantizations

Quantization Repo Size Quality
BF16 (reference) BF16 ~60 GB Reference
FP16 GGUF FP16 ~60 GB Lossless
Q8_0 GGUF Q8_0 ~32 GB Near-lossless
Q6_K GGUF Q6_K ~25 GB Excellent
Q4_K_M GGUF [You are here] ~18 GB Good

Vision (Multimodal)

This model accepts image input when paired with a vision projector (mmproj). Abliteration only modified the language backbone — the vision encoder is untouched — so the standard Meta projector works directly with this repo.

This repository bundles mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder

  • projector for Muse Glimmer 30B.

Usage (llama.cpp)

huggingface-cli download mlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF \
  --include "Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf" \
  --include "mmproj-Muse-Glimmer-30B-Q4_K_M.gguf" \
  --local-dir ./models

./build/bin/llama-mtmd-cli \
  -m ./models/Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  --mmproj ./models/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \
  --image photo.png \
  -p "Describe this image."

Ollama note: Ollama does not currently support separate mmproj files for this architecture. For image input, use llama.cpp (llama-mtmd-cli or llama-server --mmproj).

Limitations & Disclaimers

  • This is an abliterated model — it has been modified to refuse fewer prompts. Use responsibly.
  • Some refusal pathways remain (notably weapons-related content). This is not a fully uncensored model.
  • Q4_K_M quantization introduces a small quality penalty vs. higher-bit quants. For demanding tasks, consider Q6_K or Q8_0.
  • Abliteration may subtly affect output quality; (\alpha = 0.15) was chosen conservatively.
  • The vision encoder is untouched by abliteration. Image input is available via the bundled mmproj projector (llama.cpp only; see above).
  • Comply with applicable laws and regulations.

License: Apache 2.0

Changelog

v1.1.0 — vision (multimodal) support (2026-08-16)

  • Added mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder + projector, enabling image input via llama.cpp.
  • The vision tower is untouched by abliteration, so this projector matches the base model (meta-models/Muse-Glimmer-30B).
  • v1.0.0 was the initial (unversioned) text-only upload.
Downloads last month
772
GGUF
Model size
28B params
Architecture
muse-glimmer
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlasli/Muse-Glimmer-30B-Abliterated-Q4_K_M-GGUF

Quantized
(4)
this model