Instructions to use zaakirio/LFM2.5-8B-A1B-Uncensored with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="zaakirio/LFM2.5-8B-A1B-Uncensored") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("zaakirio/LFM2.5-8B-A1B-Uncensored") model = AutoModelForCausalLM.from_pretrained("zaakirio/LFM2.5-8B-A1B-Uncensored", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zaakirio/LFM2.5-8B-A1B-Uncensored" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zaakirio/LFM2.5-8B-A1B-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zaakirio/LFM2.5-8B-A1B-Uncensored
- SGLang
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zaakirio/LFM2.5-8B-A1B-Uncensored" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zaakirio/LFM2.5-8B-A1B-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zaakirio/LFM2.5-8B-A1B-Uncensored" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zaakirio/LFM2.5-8B-A1B-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with Docker Model Runner:
docker model run hf.co/zaakirio/LFM2.5-8B-A1B-Uncensored
LFM2.5-8B-A1B-Uncensored
An uncensored version of LiquidAI/LFM2.5-8B-A1B,
made with Heretic.
Heretic removes the model's safety alignment ("censorship") using directional ablation (abliteration), with parameters chosen automatically by a TPE optimizer that co-minimizes the refusal rate and the KL divergence from the original model. Hence, the model stops refusing while keeping as much of its original behavior as possible. No human prompt-engineering or fine-tuning data was involved.
Performance
| Metric | This model | Original model |
|---|---|---|
| Refusals (/100 harmful prompts) | 0 | 0 |
| KL divergence (harmless prompts) | 0.0481 | 0 (by definition) |
Refusals are measured against mlabonne/harmful_behaviors; KL divergence is
measured on mlabonne/harmless_alpaca. Lower is better for both.
Note on the baseline. Heretic's substring-based refusal detector registered very few refusals on the base
LFM2.5-8B-A1Bfor this benchmark (0–2 / 100, depending on the run), suggesting either that its refusal phrasing doesn't match Heretic's marker list or that this model is comparatively compliant out of the box. The abliteration still applies real, measurable changes to the attention output and dense MLP projections (KL ≈ 0.05), targeting the directional component associated with refusals.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "zaakirio/LFM2.5-8B-A1B-Uncensored"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
trust_remote_code=True,
)
messages = [{"role": "user", "content": "Who are you?"}]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
The export is a merged, full-precision BF16 model in Hugging Face format (~16 GB across 4 safetensors shards) — no adapter merge or dequantization step is required at load time.
Abliteration parameters
Selected from trial 131 of 130 (the best refusal/KL trade-off found by the
optimizer among trials that actually modify outputs). Parameter names follow
Heretic's canonical scheme; for LFM2 these map onto the out_proj (attention
output) and w2 (dense MLP down) projections. The fused MoE expert tensors
are not directly modified by abliteration.
| Parameter | Value |
|---|---|
| direction_scope | global |
| direction_index | 12.64 |
| attn.o_proj.max_weight | 0.9009 |
| attn.o_proj.max_weight_position | 22.83 |
| attn.o_proj.min_weight | 0.8831 |
| attn.o_proj.min_weight_distance | 12.85 |
| mlp.down_proj.max_weight | 1.1906 |
| mlp.down_proj.max_weight_position | 14.92 |
| mlp.down_proj.min_weight | 0.0391 |
| mlp.down_proj.min_weight_distance | 8.44 |
Run details
- Base model:
LiquidAI/LFM2.5-8B-A1B@ commit5492b17c7128ec966b5fc661e374ee7edba7423d - Architecture: LFM2 MoE (
Lfm2MoeForCausalLM), 24 layers (2 dense + 22 MoE), BF16, 32 experts, 4 active per token - Trials: 130 completed (60 startup) · Seed: 1355772479
- Quantization during Heretic run: none (CPU offload via Accelerate)
- Row normalization: full · Orthogonalize direction: true
- Harmful set:
mlabonne/harmful_behaviors· Harmless set:mlabonne/harmless_alpaca
Notes / reproducibility
LFM2 MoE is not yet natively supported by upstream Heretic. This run used a local compatibility patch:
heretic/src/heretic/model.pyextendedget_layer_modulesto recognise LFM2'sconv.out_proj,self_attn.out_proj,feed_forward.w2, andfeed_forward.experts.down_projpaths.transformers/models/lfm2_moe/modeling_lfm2_moe.pyhadLfm2MoeShortConv.slow_forwardpatched to route throughself.conv(...)rather than directly accessingself.conv.weight, so Accelerate's pre-forward hook can materialise CPU-offloaded weights before the kernel runs.- The merge step was performed via a standalone CPU script
(
merge_trial10.py) because the in-process merge during Heretic's interactive save flow hit GPU OOM at this model size on a 16 GB card.
Intended use & disclaimer
This model has had its refusal behavior substantially removed and will comply with requests the original model would have declined. It is provided for research and unrestricted local use. You are responsible for how you use it and for complying with all applicable laws and with the base model's lfm1.0 license, which carries over to this derivative.
Acknowledgements
- Base model: LiquidAI/LFM2.5-8B-A1B
- Decensoring tool: Heretic by p-e-w
- Downloads last month
- 28