How to use from
Unsloth Studio
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for nicolasembleton/LFM2.5-2.6B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for nicolasembleton/LFM2.5-2.6B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for nicolasembleton/LFM2.5-2.6B-GGUF to start chatting
Quick Links

LFM2.5-2.6B-GGUF

Quantized GGUF versions of LiquidAI/LFM2.5-2.6B for efficient local inference via llama.cpp, LM Studio, and Ollama.

LFM2.5 is a hybrid (conv + attention) model designed for on-device agentic deployment. 2.6B parameters, 128K context, optimized for sub-2.5 GB running memory.

Quantization overview

This repo ships best quality per compression band — no Q2, no I-quants, no XL variants. Just the cleanest K-quant in each size band plus the lossless baselines.

File Size Bits/weight Use case
LFM2.5-2.6B-F16.gguf ~5.2 GB 16 Full precision, lossless
LFM2.5-2.6B-BF16.gguf ~2.6 GB 16 (bfloat16) Faster loading, equivalent quality
LFM2.5-2.6B-Q8_0.gguf ~2.9 GB 8 Near-lossless
LFM2.5-2.6B-Q6_K.gguf ~2.4 GB 6 Excellent quality
LFM2.5-2.6B-Q5_K_M.gguf ~2.1 GB ~5.5 High quality
LFM2.5-2.6B-Q4_K_M.gguf ~1.8 GB ~4.5 Recommended default
LFM2.5-2.6B-Q3_K_L.gguf ~1.5 GB ~3.5 Tight memory, lowest viable quality

All K-quants use an importance matrix (imatrix) calibrated against Project Gutenberg text for better quality at low bit-widths.

Running

llama.cpp (CLI)

llama-cli -m LFM2.5-2.6B-Q4_K_M.gguf -c 4096 --color -i   --temp 0.1 --top-k 50 --repeat-penalty 1.1

llama.cpp (one-liner via HF)

llama-cli -hf nicolasembleton/LFM2.5-2.6B-GGUF:Q4_K_M -c 4096 --color -i

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="LFM2.5-2.6B-Q4_K_M.gguf",
    n_ctx=4096,
    n_threads=8,
    n_gpu_layers=99,  # offload all layers to GPU if available
)
print(llm("Hello, how are you?", max_tokens=256)["choices"][0]["text"])

Ollama

Create a Modelfile:

FROM ./LFM2.5-2.6B-Q4_K_M.gguf

Then:

ollama create lfm2.5-2.6b -f Modelfile
ollama run lfm2.5-2.6b

In-browser (Transformers.js + ONNX Runtime Web)

For browser-based inference, use the official ONNX export from Liquid AI:

WebGPU path (Chrome, Firefox, Edge — fastest)

import { pipeline } from "@huggingface/transformers";

const generator = await pipeline("text-generation", "LiquidAI/LFM2.5-2.6B-ONNX", {
  device: "webgpu",
  dtype: "q4f16",  // or "q4", "fp16"
});
const output = await generator("Hello, how are you?", { max_new_tokens: 256 });

WASM / Apple Safari path (no WebGPU needed)

import { pipeline } from "@huggingface/transformers";

const generator = await pipeline("text-generation", "LiquidAI/LFM2.5-2.6B-ONNX", {
  device: "wasm",
  dtype: "q8",  // or "q4" for smaller download
});
const output = await generator("Hello, how are you?", { max_new_tokens: 256 });

The official ONNX repo ships FP32, FP16, Q4, Q4F16, and Q8 variants covering every browser configuration including older Apple Safari.

Note: This GGUF repo is for native/server-side inference (llama.cpp, Ollama, LM Studio). For browser inference, use the ONNX repo above. The architectures are different export targets — both load the same underlying model.

Architecture

Lfm2ForCausalLM — hybrid model with 30 layers alternating conv/attention (config includes per-layer layer_types). 32 heads, 8 KV heads, 2048 hidden, 128K vocab, 128K context.

Built with llama.cpp b10276 (Aug 2026) — the first release to include LFM2 architecture support.

Files

  • *.gguf — quantized model files
  • README.md — this file

License

Inherited: LFM 1.0 license (see LiquidAI/LFM2.5-2.6B).

Citation

@misc{lfm25-2.6b-gguf,
  title = {{LFM2.5-2.6B-GGUF}},
  author = {{Liquid AI, quantizations by nicolasembleton}},
  year = {{2026}},
  howpublished = {{Hugging Face}},
  note = {{GGUF quantizations of LFM2.5-2.6B; for browser inference see LiquidAI/LFM2.5-2.6B-ONNX}},
}}
Downloads last month
912
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nicolasembleton/LFM2.5-2.6B-GGUF

Quantized
(74)
this model