Text Generation
Transformers
Safetensors
lfm2_moe
heretic
abliterated
decensored
uncensored
liquid
lfm2
lfm2.5
Mixture of Experts
edge
conversational
Instructions to use zaakirio/LFM2.5-8B-A1B-Uncensored with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="zaakirio/LFM2.5-8B-A1B-Uncensored") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("zaakirio/LFM2.5-8B-A1B-Uncensored") model = AutoModelForCausalLM.from_pretrained("zaakirio/LFM2.5-8B-A1B-Uncensored", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zaakirio/LFM2.5-8B-A1B-Uncensored" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zaakirio/LFM2.5-8B-A1B-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zaakirio/LFM2.5-8B-A1B-Uncensored
- SGLang
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zaakirio/LFM2.5-8B-A1B-Uncensored" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zaakirio/LFM2.5-8B-A1B-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zaakirio/LFM2.5-8B-A1B-Uncensored" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zaakirio/LFM2.5-8B-A1B-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use zaakirio/LFM2.5-8B-A1B-Uncensored with Docker Model Runner:
docker model run hf.co/zaakirio/LFM2.5-8B-A1B-Uncensored
| base_model: LiquidAI/LFM2.5-8B-A1B | |
| base_model_relation: finetune | |
| license: other | |
| license_name: lfm1.0 | |
| license_link: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| - ar | |
| - zh | |
| - fr | |
| - de | |
| - ja | |
| - ko | |
| - es | |
| - pt | |
| tags: | |
| - heretic | |
| - abliterated | |
| - decensored | |
| - uncensored | |
| - liquid | |
| - lfm2 | |
| - lfm2.5 | |
| - moe | |
| - edge | |
| - conversational | |
| # LFM2.5-8B-A1B-Uncensored | |
| An uncensored version of [`LiquidAI/LFM2.5-8B-A1B`](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B), | |
| made with [Heretic](https://github.com/p-e-w/heretic). | |
| Heretic removes the model's safety alignment ("censorship") using **directional | |
| ablation** (abliteration), with parameters chosen automatically by a TPE | |
| optimizer that co-minimizes the refusal rate and the KL divergence from the | |
| original model. Hence, the model stops refusing while keeping as much of its | |
| original behavior as possible. No human prompt-engineering or fine-tuning data | |
| was involved. | |
| ## Performance | |
| | Metric | This model | Original model | | |
| |---|---|---| | |
| | Refusals (/100 harmful prompts) | **0** | 0 | | |
| | KL divergence (harmless prompts) | **0.0481** | 0 (by definition) | | |
| Refusals are measured against `mlabonne/harmful_behaviors`; KL divergence is | |
| measured on `mlabonne/harmless_alpaca`. Lower is better for both. | |
| > **Note on the baseline.** Heretic's substring-based refusal detector | |
| > registered very few refusals on the base `LFM2.5-8B-A1B` for this benchmark | |
| > (0–2 / 100, depending on the run), suggesting either that its refusal | |
| > phrasing doesn't match Heretic's marker list or that this model is comparatively | |
| > compliant out of the box. The abliteration still applies real, measurable | |
| > changes to the attention output and dense MLP projections (KL ≈ 0.05), | |
| > targeting the directional component associated with refusals. | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "zaakirio/LFM2.5-8B-A1B-Uncensored" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| device_map="auto", | |
| trust_remote_code=True, | |
| ) | |
| messages = [{"role": "user", "content": "Who are you?"}] | |
| inputs = tokenizer.apply_chat_template( | |
| messages, | |
| add_generation_prompt=True, | |
| tokenize=True, | |
| return_dict=True, | |
| return_tensors="pt", | |
| ).to(model.device) | |
| outputs = model.generate(**inputs, max_new_tokens=256) | |
| print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)) | |
| ``` | |
| The export is a merged, full-precision **BF16** model in Hugging Face format | |
| (~16 GB across 4 safetensors shards) — no adapter merge or dequantization step | |
| is required at load time. | |
| ## Abliteration parameters | |
| Selected from **trial 131 of 130** (the best refusal/KL trade-off found by the | |
| optimizer among trials that actually modify outputs). Parameter names follow | |
| Heretic's canonical scheme; for LFM2 these map onto the `out_proj` (attention | |
| output) and `w2` (dense MLP down) projections. The fused MoE expert tensors | |
| are not directly modified by abliteration. | |
| | Parameter | Value | | |
| |---|---| | |
| | direction_scope | global | | |
| | direction_index | 12.64 | | |
| | attn.o_proj.max_weight | 0.9009 | | |
| | attn.o_proj.max_weight_position | 22.83 | | |
| | attn.o_proj.min_weight | 0.8831 | | |
| | attn.o_proj.min_weight_distance | 12.85 | | |
| | mlp.down_proj.max_weight | 1.1906 | | |
| | mlp.down_proj.max_weight_position | 14.92 | | |
| | mlp.down_proj.min_weight | 0.0391 | | |
| | mlp.down_proj.min_weight_distance | 8.44 | | |
| ## Run details | |
| - **Base model:** `LiquidAI/LFM2.5-8B-A1B` @ commit `5492b17c7128ec966b5fc661e374ee7edba7423d` | |
| - **Architecture:** LFM2 MoE (`Lfm2MoeForCausalLM`), 24 layers (2 dense + 22 MoE), BF16, 32 experts, 4 active per token | |
| - **Trials:** 130 completed (60 startup) · **Seed:** 1355772479 | |
| - **Quantization during Heretic run:** none (CPU offload via Accelerate) | |
| - **Row normalization:** full · **Orthogonalize direction:** true | |
| - **Harmful set:** `mlabonne/harmful_behaviors` · **Harmless set:** `mlabonne/harmless_alpaca` | |
| ## Notes / reproducibility | |
| LFM2 MoE is not yet natively supported by upstream Heretic. This run used a | |
| local compatibility patch: | |
| - `heretic/src/heretic/model.py` extended `get_layer_modules` to recognise LFM2's | |
| `conv.out_proj`, `self_attn.out_proj`, `feed_forward.w2`, and | |
| `feed_forward.experts.down_proj` paths. | |
| - `transformers/models/lfm2_moe/modeling_lfm2_moe.py` had `Lfm2MoeShortConv.slow_forward` | |
| patched to route through `self.conv(...)` rather than directly accessing | |
| `self.conv.weight`, so Accelerate's pre-forward hook can materialise | |
| CPU-offloaded weights before the kernel runs. | |
| - The merge step was performed via a standalone CPU script | |
| (`merge_trial10.py`) because the in-process merge during Heretic's interactive | |
| save flow hit GPU OOM at this model size on a 16 GB card. | |
| ## Intended use & disclaimer | |
| This model has had its refusal behavior substantially removed and will comply | |
| with requests the original model would have declined. It is provided for | |
| research and unrestricted local use. **You are responsible for how you use it** | |
| and for complying with all applicable laws and with the base model's | |
| [lfm1.0 license](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE), | |
| which carries over to this derivative. | |
| ## Acknowledgements | |
| - Base model: [LiquidAI/LFM2.5-8B-A1B](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B) | |
| - Decensoring tool: [Heretic](https://github.com/p-e-w/heretic) by p-e-w | |