LFM2.5 Encoder 230M + SigLIP2

A public GPTQ INT4 release of an encoder-only image-text model built by augmenting Liquid AI's LFM2.5 Encoder 230M with the google/siglip2-base-patch16-256 vision encoder and a learned 32-token visual projector.

This is an external vision augmentation of LFM2.5. The original LFM2.5 checkpoint is not claimed to be natively multimodal.

What it does

The model produces image-text representations and image-text matching scores for:

  • Image-to-text retrieval
  • Text-to-image retrieval
  • Image-text matching
  • Representation extraction
  • Encoder-side masked-token experiments

It is not an autoregressive model and does not generate captions or chat responses.

Architecture

Image
  ↓
SigLIP2 base patch16 256
  [batch, 256, 768]
  ↓
Learned visual projector / 32-token adapter
  [batch, 32, 1024]
  ↓
LFM2.5 bidirectional encoder
  [batch, sequence, 1024]
  ↓
Image-text embedding or matching score

The image branch uses BOS, an image marker, and projected visual tokens. During clean retrieval evaluation, the candidate caption is never provided to the image branch.

Quick start

Clone the repository and install the tested runtime dependencies:

git clone https://huggingface.co/konic-labs/LFM2.5-multimodal-encoder-230M
cd LFM2.5-multimodal-encoder-230M

pip install torch transformers pillow accelerate safetensors compressed-tensors

GPTQ loading requires a compressed-tensors installation compatible with the model configuration. The tested quantization details are recorded in recipe.yaml and compression_manifest.json.

The SigLIP2 vision model is downloaded from Hugging Face on first use. Run one image-text matching example:

python run_multimodal.py \
  --model . \
  --image /path/to/image.jpg \
  --text "A photograph of a small red vehicle beside a building." \
  --device cuda:0 \
  --dtype float32

The runner reports the visual feature shape, projected token shape, fused hidden-state shape, matching score, and finite-output status.

Expected shapes include:

visual_features: [1, 256, 768]
image_tokens:    [1, 32, 1024]
hidden_state:    [1, sequence, 1024]

The task adapter contains the matching head and the projector state used by the bundled runner. projector.pt is also provided as a standalone projector artifact for modular experiments.

Main files

File Purpose
model.safetensors GPTQ INT4 LFM2.5 encoder weights
config.json Transformers model configuration
modeling_lfm2_bidirectional.py Bidirectional LFM2 runtime implementation
multimodal_task_adapter.pt Projector and image-text matching head
projector.pt Standalone SigLIP2-to-LFM projector
multimodal_reference_manifest.json Vision and adapter configuration
compression_manifest.json GPTQ method and calibration provenance
recipe.yaml Quantization recipe
run_multimodal.py Local image-text inference example
evaluation/ Saved metrics and deep-evaluation summaries

Evaluation

All clean evaluations use an image-only representation for retrieval. Retrieval metrics were computed on 12,500 held-out MONET pairs with 200 bootstrap resamples for confidence intervals.

MONET test set

Direction R@1 R@5 R@10 Mean rank
Image → text 0.1091 0.2959 0.4090 63.38
Text → image 0.1460 0.3418 0.4590 53.78

External CC3M subset

A bounded, 2,000-example CC3M subset was used for external evaluation:

Direction R@1 R@5 R@10 Mean rank
Image → text 0.0865 0.2190 0.3060 137.34
Text → image 0.1145 0.2835 0.3865 112.76

Visual ablation

On 512 MONET examples:

Visual input Image → text R@1 Text → image R@1
Trained visual tokens 0.4883 0.5352
Zero visual tokens 0.0020 0.0000
Random visual tokens 0.0020 0.0000
Shuffled visual tokens 0.0117 0.0215

These controls show that the learned visual tokens materially affect the image representation.

Matching

Cyclic-negative matching AUROC is 0.9733 on MONET. Text-nearest hard-negative AUROC is 0.4961, so the current model is substantially better at coarse compatibility than fine-grained semantic discrimination.

Compression details and size comparison

The released LFM2.5 body uses GPTQ INT4 W4A16 compression:

Compression setting Value
Method GPTQ INT4
Weight/activation format W4A16
Quantized layers Linear layers; lm_head excluded
Block size 128
Activation order Static
Dampening fraction 0.01
Calibration 256 fused multimodal inputs_embeds samples
Maximum calibration sequence length 160 tokens
Compressor revision 8cec0acc1931de6f8f73257151ab7007c14dbf4e

The calibration samples contained actual projected visual tokens rather than text-only inputs, so the quantization pass reflects the multimodal path.

Storage comparison

Artifact Model weights Full package Relative to original package
Original LFM2.5 Encoder 230M FP32 918.79 MB 923.65 MB 100%
Clean BF16 reference 459.40 MB 474.81 MB 51.4%
Clean GPTQ INT4 release 355.03 MB 370.46 MB 40.1%

The GPTQ release is approximately 59.9% smaller than the original package and 22.0% smaller than the clean BF16 package. The external SigLIP2 vision tower is not included in these sizes and is downloaded separately.

Fidelity and runtime

Metric GPTQ INT4
Hidden cosine versus BF16 0.9427
Logit cosine versus BF16 0.9654
Multimodal latency, batch 4 45.7 ms on NVIDIA L4
Peak multimodal VRAM 2.38 GiB on NVIDIA L4

The current GPTQ backend is intended primarily for weight-size reduction. Native BF16 is faster and uses less runtime VRAM in the tested environment.

Limitations

  • Encoder-only: no generation is supported.
  • Hard-negative matching remains near chance.
  • Masked-token prediction does not currently improve over text-only or zero-image controls.
  • External evaluation is a bounded CC3M subset rather than a full benchmark.
  • The SigLIP2 vision tower remains an external Hugging Face dependency.

Reproducibility

The evaluation/ directory contains the saved BF16, GPTQ, and external CC3M metrics, retrieval summary, matching results, ablations, robustness results, and the detailed deep-evaluation summary.

The compact release corresponds to the clean image-only training branch and should be used instead of the earlier paired-caption artifact.

Downloads last month
27
Safetensors
Model size
0.2B params
Tensor type
F32
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including konic-labs/LFM2.5-multimodal-encoder-230M