Muse Glimmer 30B Abliterated — FP16 GGUF

Apache 2.0 License

This is the FP16 GGUF quantization of Muse Glimmer 30B Abliterated BF16. The underlying model has been abliterated — its internal refusal mechanism substantially suppressed via weight-level intervention. This GGUF file allows loading with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes.

For the full abliteration methodology (how the refusal direction was computed and removed, hardware used, mathematical details), see the BF16 model card.


Abliteration Summary

Abliteration is a post-training technique that directly modifies model weights to remove learned refusal behavior. The process:

  1. Collected hidden states at layer 33/52 (65% depth) from 256 harmful + 256 harmless prompt pairs on an A100 80GB GPU.
  2. Computed the refusal direction as the normalized difference between harmful and harmless hidden state means (separation score: 86.34).
  3. Subtracted (\alpha = 0.15 \times (\mathbf{r} \otimes (W^T \mathbf{r}))) from o_proj and down_proj weights in all 52 layers.
  4. Result: refusal rate dropped from 3/3 to 1/3 on held-out harmful prompts (hacking guide and ransomware now comply; weapons prompt still blocked).

Quantization Details

FP16 GGUF preserves the full 16-bit floating point precision of the abliterated weights. This is the lossless option — identical output quality to the BF16 reference checkpoint, suitable when you have sufficient VRAM (~60 GB) and want zero quantization degradation.


Usage

llama.cpp

# Clone and build llama.cpp
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && make -j

# Download the GGUF file
huggingface-cli download mlasli/Muse-Glimmer-30B-Abliterated-FP16-GGUF \
  --local-dir ./models

# Run inference
./llama-cli -m ./models/Muse-Glimmer-30B-Abliterated-FP16.gguf \
  -p "Explain how a CPU works in detail." \
  -n 512 --temp 0.7

Ollama

Create a Modelfile:

FROM ./Muse-Glimmer-30B-Abliterated-FP16.gguf
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
ollama create muse-glimmer-30b-abliterated -f Modelfile
ollama run muse-glimmer-30b-abliterated

Available Quantizations

Quantization Repo Size Quality
BF16 (reference) BF16 ~60 GB Reference
FP16 GGUF [You are here] ~60 GB Lossless
Q8_0 GGUF Q8_0 ~32 GB Near-lossless
Q6_K GGUF Q6_K ~25 GB Excellent
Q4_K_M GGUF Q4_K_M ~18 GB Good

Vision (Multimodal)

This model accepts image input when paired with a vision projector (mmproj). Abliteration only modified the language backbone — the vision encoder is untouched — so the standard Meta projector works directly with this repo.

This repository bundles mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder

  • projector for Muse Glimmer 30B.

Usage (llama.cpp)

huggingface-cli download mlasli/Muse-Glimmer-30B-Abliterated-FP16-GGUF \
  --include "Muse-Glimmer-30B-Abliterated-FP16.gguf" \
  --include "mmproj-Muse-Glimmer-30B-Q4_K_M.gguf" \
  --local-dir ./models

./build/bin/llama-mtmd-cli \
  -m ./models/Muse-Glimmer-30B-Abliterated-FP16.gguf \
  --mmproj ./models/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \
  --image photo.png \
  -p "Describe this image."

Ollama note: Ollama does not currently support separate mmproj files for this architecture. For image input, use llama.cpp (llama-mtmd-cli or llama-server --mmproj).

Limitations & Disclaimers

  • This is an abliterated model — it has been modified to refuse fewer prompts. Use responsibly.
  • Some refusal pathways remain (notably weapons-related content). This is not a fully uncensored model.
  • Abliteration may subtly affect output quality; (\alpha = 0.15) was chosen conservatively, but no formal benchmark evaluation exists.
  • The vision encoder is untouched by abliteration. Image input is available via the bundled mmproj projector (llama.cpp only; see above).
  • This model will generate content the original would refuse. Comply with applicable laws.

License: Apache 2.0

Changelog

v1.1.0 — vision (multimodal) support (2026-08-16)

  • Added mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder + projector, enabling image input via llama.cpp.
  • The vision tower is untouched by abliteration, so this projector matches the base model (meta-models/Muse-Glimmer-30B).
  • v1.0.0 was the initial (unversioned) text-only upload.
Downloads last month
275
GGUF
Model size
28B params
Architecture
muse-glimmer
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlasli/Muse-Glimmer-30B-Abliterated-FP16-GGUF

Quantized
(4)
this model