How to use from
Pi
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf WhiskyAKM/LFM2.5-2.6B-GGUF:
Configure the model in Pi
# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "WhiskyAKM/LFM2.5-2.6B-GGUF:"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

LFM2.5-2.6B GGUF

GGUF quantized versions of LiquidAI/LFM2.5-2.6B, a high-performance hybrid model designed for on-device deployment, featuring a 128K context window and advanced agentic capabilities.

Model Overview

LFM2.5-2.6B is part of the LFM2.5 family, building on the LFM2 architecture to provide best-in-class performance for its size. It is specifically optimized for agentic workloads, tool use, and long-context workflows, offering competitive performance against models 4x its size.

Key features include:

  • Agentic Post-Training: Trained using agentic reinforcement learning for improved tool use and instruction following.
  • Efficient Inference: Designed for high-speed execution on both CPU and GPU.
  • Reasoning Capabilities: A pure reasoning model that utilizes a <think> tag to reason before answering.
  • Massive Context: Supports up to 131,072 tokens.

Model Architecture

Property Value
Architecture LFM2
Parameters 2.69B
Layers 30 (22 conv + 8 GQA)
Context Length 131,072
Vocabulary Size 128,000
Training Budget 34 Trillion Tokens
Supported Languages English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish

Available GGUF Files

File Quantization Use Case
lfm2.5-2.6b.gguf FP16/BF16 Max precision, reference model
lfm2.5-2.6b-Q8_0.gguf Q8_0 Near-lossless, high fidelity
lfm2.5-2.6b-Q6_K.gguf Q6_K Very high quality, recommended for quality
lfm2.5-2.6b-Q5_K_M.gguf Q5_K_M High quality, balanced
lfm2.5-2.6b-Q5_K_S.gguf Q5_K_S High quality, slightly smaller
lfm2.5-2.6b-Q4_K_M.gguf Q4_K_M Good quality, recommended default
lfm2.5-2.6b-Q4_K_S.gguf Q4_K_S Smaller, acceptable quality
lfm2.5-2.6b-Q4_0.gguf Q4_0 Legacy quant, fastest inference

Recommended: Q4_K_M or Q5_K_M offer the best quality-to-size trade-off for most use cases.

Usage

llama.cpp CLI

./llama-cli \
  -m lfm2.5-2.6b-Q4_K_M.gguf \
  -p "What is the capital of France?" \
  --temp 0.1 --top-k 50 --repeat-penalty 1.1

llama-server (OpenAI-compatible API)

./llama-server \
  -m lfm2.5-2.6b-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080

Chat Template & Reasoning

LFM2.5 uses a ChatML-like format. It is a reasoning model that automatically adds a <think> tag when starting an assistant answer to process its logic before providing the final response.

Example format:

<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
<think>
... reasoning process ...
</think>
C. elegans is a species of small roundworm...<|im_end|>

Tool Calling

LFM2.5 supports Pythonic function calling. It outputs function calls between <|tool_call_start|> and <|tool_call_end|> tokens.

Generation Parameters

Recommended parameters for optimal performance:

Parameter Value
Temperature 0.1
Top-K 50
Repetition Penalty 1.1

Quantization

These GGUF files were created using llama.cpp tools to enable efficient local deployment on CPUs and GPUs with reduced memory footprints.

Acknowledgements

License

LFM 1.0 License

Downloads last month
-
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WhiskyAKM/LFM2.5-2.6B-GGUF

Quantized
(9)
this model

Collection including WhiskyAKM/LFM2.5-2.6B-GGUF