FadedRedStar commited on
Commit
16edadc
ยท
verified ยท
1 Parent(s): e4f30ab

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +149 -16
README.md CHANGED
@@ -2,35 +2,168 @@
2
  base_model: coder3101/LFM2.5-350M-heretic
3
  base_model_relation: quantized
4
  library_name: gguf
5
- license: apache-2.0
 
 
6
  language:
7
- - multilingual
 
 
 
 
 
 
 
 
8
  pipeline_tag: text-generation
9
  tags:
10
  - gguf
 
11
  - text-generation
12
- - reasoning
13
- - thinking
14
- - agentic
15
- - tool-use
16
- - abliterated
17
  - uncensored
18
- - lfm2
19
  - imatrix
 
20
  - conversational
 
 
 
 
21
  ---
22
 
23
- # LFM2.5-350M-heretic-imatrix โ€” GGUF
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24
 
25
  > [!NOTE]
26
- > This card is a work in progress. Full architecture details, abliteration parameters, and inference instructions will be added shortly.
27
 
28
- This repository hosts importance-matrix (imatrix) calibrated GGUF weights for **LFM2.5-350M-heretic**, quantized from the source floating-point tensors provided by [coder3101/LFM2.5-350M-heretic](https://huggingface.co/coder3101/LFM2.5-350M-heretic).
 
 
 
29
 
30
- ## Repository Files
31
 
32
- | Filename | Format | Size | llama.cpp Build | Description |
 
 
 
 
 
 
 
 
33
  |---|---|---|---|---|
34
- | `LFM2.5-350M-heretic-IQ4_NL-imatrix.gguf` | IQ4_NL | 219 MB | b9860 | Main model weights, imatrix-calibrated, smaller non-linear 4-bit quant |
35
- | `LFM2.5-350M-heretic-Q4_K_M-imatrix.gguf` | Q4_K_M | 229 MB | b9860 | Main model weights, imatrix-calibrated |
36
- | `LFM2.5-350M-heretic-Q5_K_M-imatrix.gguf` | Q5_K_M | 260 MB | b9860 | Main model weights, imatrix-calibrated, higher fidelity |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  base_model: coder3101/LFM2.5-350M-heretic
3
  base_model_relation: quantized
4
  library_name: gguf
5
+ license: other
6
+ license_name: lfm-1.0
7
+ license_link: https://huggingface.co/LiquidAI/LFM2.5-350M/blob/main/LICENSE
8
  language:
9
+ - en
10
+ - ar
11
+ - zh
12
+ - fr
13
+ - de
14
+ - ja
15
+ - ko
16
+ - es
17
+ - pt
18
  pipeline_tag: text-generation
19
  tags:
20
  - gguf
21
+ - llama.cpp
22
  - text-generation
23
+ - heretic
24
+ - liquid-ai
 
 
 
25
  - uncensored
 
26
  - imatrix
27
+ - abliterated
28
  - conversational
29
+ - iq4_nl
30
+ - q4_k_m
31
+ - q5_k_m
32
+ quantized_by: FadedRedStar
33
  ---
34
 
35
+ # ๐Ÿค– LFM2.5-350M-heretic โ€” Importance Matrix GGUF
36
+
37
+ This repository hosts importance-matrix (imatrix) optimized GGUF weights, available in multiple quantization formats, for **LFM2.5-350M-heretic**, quantized from the source floating-point tensors provided by [coder3101/LFM2.5-350M-heretic](https://huggingface.co/coder3101/LFM2.5-350M-heretic).
38
+
39
+ **๐Ÿ”„ Sister Repository:** Check out the [Standard GGUF Sister Repository](https://huggingface.co/FadedRedStar/LFM2.5-350M-heretic-GGUF) for uncalibrated and full 8-bit precision options.
40
+
41
+ ## ๐ŸŽฏ Matrix-Weighted Calibration (Imatrix)
42
+
43
+ An **Importance Matrix (imatrix)** calculation tracks activations across network layers using a calibration sequence, then weights the quantization process to preserve the parameters that matter most for output quality โ€” improving fidelity at low bit depths.
44
+
45
+ โžก๏ธ **Calibration dataset:** Bartowski's `calibration_datav5.txt`.
46
+
47
+ > [!NOTE]
48
+ > * **`IQ4_NL` is included** because the matrix enables a non-linear 4-bit format that outperforms standard linear 4-bit quantization.
49
+ > * **`Q8_0` is absent** because 8-bit quantization already introduces near-zero degradation, making calibration unnecessary โ€” see the standard sister repository for that variant.
50
+
51
+ ## โ„น๏ธ Model Profile & Core Features
52
+
53
+ **LFM2.5-350M** is the smallest text-only model in Liquid AI's **Liquid Foundation Model 2.5** series, built for extreme on-device and edge deployment. It shares the family's hybrid architecture of double-gated LIV (Liquid, Input-adaptive, Value-selective) convolution layers interleaved with GQA (Grouped Query Attention) layers, pre-trained on 28 trillion tokens with large-scale reinforcement learning post-training. Despite its size, it is tuned for **instruction following, lightweight tool calling, and structured data extraction**, with day-one support across llama.cpp, MLX, vLLM, SGLang, ONNX, and OpenVINO.
54
+
55
+ The **heretic** suffix denotes post-processing via the **[Heretic v1.3.0](https://github.com/p-e-w/heretic)** method performed by [coder3101](https://huggingface.co/coder3101), which removes refusal conditioning while preserving the model's lightweight instruction-following behavior.
56
+
57
+ ## ๐Ÿ“‹ Technical Specifications
58
+
59
+ | Property | Value |
60
+ |---|---|
61
+ | **Base Architecture** | LFM2 hybrid (double-gated LIV conv + GQA) |
62
+ | **Developed by** | Liquid AI |
63
+ | **Total Parameters** | 350M |
64
+ | **Primary Use** | Instruction following, lightweight tool calling, structured extraction |
65
+ | **Context Window** | 131,072 tokens |
66
+ | **Training Budget** | 28 trillion tokens |
67
+ | **Languages** | English, Arabic, Chinese, French, German, Japanese, Korean, Spanish, Portuguese |
68
+ | **Abliteration Tool** | Heretic v1.3.0 |
69
+ | **Prompt Format** | ChatML |
70
+
71
+ ## ๐Ÿ› ๏ธ Heretic Overrides (ARA)
72
+
73
+ | Property | Value |
74
+ |---|---|
75
+ | **direction_index** | per layer |
76
+ | **attn.o_proj.max_weight** | 1.08 |
77
+ | **attn.o_proj.max_weight_position** | 10.46 |
78
+ | **attn.o_proj.min_weight** | 0.87 |
79
+ | **attn.o_proj.min_weight_distance** | 3.56 |
80
+ | **mlp.down_proj.max_weight** | 1.44 |
81
+ | **mlp.down_proj.max_weight_position** | 12.00 |
82
+ | **mlp.down_proj.min_weight** | 1.22 |
83
+ | **mlp.down_proj.min_weight_distance** | 1.97 |
84
+
85
+ ## ๐Ÿ“Š Refusal Bypass Metrics
86
 
87
  > [!NOTE]
88
+ > The metrics below are self-reported by the original model author ([coder3101](https://huggingface.co/coder3101)) and have not been independently reproduced.
89
 
90
+ | Metric | This model | Original ([LiquidAI/LFM2.5-350M](https://huggingface.co/LiquidAI/LFM2.5-350M)) |
91
+ |---|---|---|
92
+ | **KL divergence** | 0.0440 | 0 *(by definition)* |
93
+ | **Refusals** | 6/100 | 90/100 |
94
 
95
+ ## ๐Ÿงฎ Numerical & Tensor Formats
96
 
97
+ | Property | Value |
98
+ |---|---|
99
+ | **Quantization Types** | IQ4_NL, Q4_K_M, Q5_K_M (all with imatrix calibration) |
100
+ | **Importance Matrix** | Bartowski's `calibration_datav5.txt` |
101
+
102
+ ## ๐Ÿ“ฆ Available Model Files
103
+
104
+ **Main model weights**
105
+ | Filename | Quantization | llama.cpp Build | Size | Download |
106
  |---|---|---|---|---|
107
+ | `LFM2.5-350M-heretic-IQ4_NL-imatrix.gguf` | `IQ4_NL` | `b9860` | 209 MB | [๐Ÿ“ฅ Download](https://huggingface.co/FadedRedStar/LFM2.5-350M-heretic-imatrix-GGUF/resolve/main/LFM2.5-350M-heretic-IQ4_NL-imatrix.gguf) |
108
+ | `LFM2.5-350M-heretic-Q4_K_M-imatrix.gguf` | `Q4_K_M` | `b9860` | 219 MB | [๐Ÿ“ฅ Download](https://huggingface.co/FadedRedStar/LFM2.5-350M-heretic-imatrix-GGUF/resolve/main/LFM2.5-350M-heretic-Q4_K_M-imatrix.gguf) |
109
+ | `LFM2.5-350M-heretic-Q5_K_M-imatrix.gguf` | `Q5_K_M` | `b9860` | 248 MB | [๐Ÿ“ฅ Download](https://huggingface.co/FadedRedStar/LFM2.5-350M-heretic-imatrix-GGUF/resolve/main/LFM2.5-350M-heretic-Q5_K_M-imatrix.gguf) |
110
+
111
+ ## ๐ŸŽ›๏ธ Component Pairing Guide
112
+
113
+ Download exactly **one** main weights file:
114
+
115
+ * **`IQ4_NL`**: Non-linear 4-bit format, best choice for constrained memory when imatrix calibration is present.
116
+ * **`Q4_K_M`**: Balanced 4-bit format suitable for most everyday use.
117
+ * **`Q5_K_M`**: Higher-fidelity mid-range format recommended as a general default.
118
+
119
+ ## โšก Deployment & Execution Commands
120
+
121
+ > [!NOTE]
122
+ > Liquid AI recommends the following generation parameters for best results: `temperature: 0.1`, `top_k: 50`, `repetition_penalty: 1.05`.
123
+
124
+ > [!TIP]
125
+ > Swap the `-m` filename below for either quantized file depending on your size/quality trade-off preference.
126
+
127
+ ### llama.cpp CLI
128
+
129
+ ```bash
130
+ ./llama-cli \
131
+ -m LFM2.5-350M-heretic-IQ4_NL-imatrix.gguf \
132
+ -c 8192 \
133
+ -ngl 99 \
134
+ --temp 0.3 \
135
+ --top-k 40 \
136
+ --repeat-penalty 1.05 \
137
+ -p "<|im_start|>system\nYou are a concise, helpful assistant.<|im_end|>\n<|im_start|>user\nState the capital of Italy and one interesting fact about it.<|im_end|>\n<|im_start|>assistant\n"
138
+ ```
139
+
140
+ ### OpenAI-Compatible API Server
141
+
142
+ ```bash
143
+ ./llama-server \
144
+ --host 0.0.0.0 \
145
+ --port 8080 \
146
+ -m LFM2.5-350M-heretic-IQ4_NL-imatrix.gguf \
147
+ -c 16384 \
148
+ -ngl 99 \
149
+ --flash-attn
150
+ ```
151
+
152
+ ## ๐Ÿ’ฌ Chat Templates & Prompt Design (ChatML)
153
+
154
+ ```text
155
+ <|im_start|>system
156
+ You are a capable assistant. Follow instructions precisely.<|im_end|>
157
+ <|im_start|>user
158
+ Your task or query here.<|im_end|>
159
+ <|im_start|>assistant
160
+ ```
161
+
162
+ ## โš ๏ธ Safety & Operational Notes
163
+
164
+ - This model is abliterated and will generate content that standard aligned models refuse. Use responsibly and in compliance with applicable laws.
165
+ - This is a **text-only** model โ€” it has no vision encoder and cannot process images.
166
+ - Despite its small footprint, IFBench and structured-extraction benchmarks show substantial generational gains over LFM2 predecessors.
167
+ - Best suited for constrained hardware: CPUs, NPUs, and edge devices rather than complex reasoning workloads.
168
+ - Imatrix calibration improves perplexity recovery compared to non-imatrix quantization, particularly on low-frequency tokens.
169
+ - IQ4_NL produces a smaller file than Q4_K_M and tends to run faster on CPU and ARM devices; imatrix calibration narrows the quality gap between the two formats considerably.