Instructions to use konic-labs/LFM2.5-multimodal-encoder-230M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use konic-labs/LFM2.5-multimodal-encoder-230M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="konic-labs/LFM2.5-multimodal-encoder-230M", trust_remote_code=True)# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("konic-labs/LFM2.5-multimodal-encoder-230M", trust_remote_code=True) model = AutoModelForMaskedLM.from_pretrained("konic-labs/LFM2.5-multimodal-encoder-230M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
LFM2.5 Encoder 230M + SigLIP2
A public GPTQ INT4 release of an encoder-only image-text model built by augmenting
Liquid AI's LFM2.5 Encoder 230M with the google/siglip2-base-patch16-256 vision
encoder and a learned 32-token visual projector.
This is an external vision augmentation of LFM2.5. The original LFM2.5 checkpoint is not claimed to be natively multimodal.
What it does
The model produces image-text representations and image-text matching scores for:
- Image-to-text retrieval
- Text-to-image retrieval
- Image-text matching
- Representation extraction
- Encoder-side masked-token experiments
It is not an autoregressive model and does not generate captions or chat responses.
Architecture
Image
↓
SigLIP2 base patch16 256
[batch, 256, 768]
↓
Learned visual projector / 32-token adapter
[batch, 32, 1024]
↓
LFM2.5 bidirectional encoder
[batch, sequence, 1024]
↓
Image-text embedding or matching score
The image branch uses BOS, an image marker, and projected visual tokens. During clean retrieval evaluation, the candidate caption is never provided to the image branch.
Quick start
Clone the repository and install the tested runtime dependencies:
git clone https://huggingface.co/konic-labs/LFM2.5-multimodal-encoder-230M
cd LFM2.5-multimodal-encoder-230M
pip install torch transformers pillow accelerate safetensors compressed-tensors
GPTQ loading requires a compressed-tensors installation compatible with the model
configuration. The tested quantization details are recorded in recipe.yaml and
compression_manifest.json.
The SigLIP2 vision model is downloaded from Hugging Face on first use. Run one image-text matching example:
python run_multimodal.py \
--model . \
--image /path/to/image.jpg \
--text "A photograph of a small red vehicle beside a building." \
--device cuda:0 \
--dtype float32
The runner reports the visual feature shape, projected token shape, fused hidden-state shape, matching score, and finite-output status.
Expected shapes include:
visual_features: [1, 256, 768]
image_tokens: [1, 32, 1024]
hidden_state: [1, sequence, 1024]
The task adapter contains the matching head and the projector state used by the bundled
runner. projector.pt is also provided as a standalone projector artifact for modular
experiments.
Main files
| File | Purpose |
|---|---|
model.safetensors |
GPTQ INT4 LFM2.5 encoder weights |
config.json |
Transformers model configuration |
modeling_lfm2_bidirectional.py |
Bidirectional LFM2 runtime implementation |
multimodal_task_adapter.pt |
Projector and image-text matching head |
projector.pt |
Standalone SigLIP2-to-LFM projector |
multimodal_reference_manifest.json |
Vision and adapter configuration |
compression_manifest.json |
GPTQ method and calibration provenance |
recipe.yaml |
Quantization recipe |
run_multimodal.py |
Local image-text inference example |
evaluation/ |
Saved metrics and deep-evaluation summaries |
Evaluation
All clean evaluations use an image-only representation for retrieval. Retrieval metrics were computed on 12,500 held-out MONET pairs with 200 bootstrap resamples for confidence intervals.
MONET test set
| Direction | R@1 | R@5 | R@10 | Mean rank |
|---|---|---|---|---|
| Image → text | 0.1091 | 0.2959 | 0.4090 | 63.38 |
| Text → image | 0.1460 | 0.3418 | 0.4590 | 53.78 |
External CC3M subset
A bounded, 2,000-example CC3M subset was used for external evaluation:
| Direction | R@1 | R@5 | R@10 | Mean rank |
|---|---|---|---|---|
| Image → text | 0.0865 | 0.2190 | 0.3060 | 137.34 |
| Text → image | 0.1145 | 0.2835 | 0.3865 | 112.76 |
Visual ablation
On 512 MONET examples:
| Visual input | Image → text R@1 | Text → image R@1 |
|---|---|---|
| Trained visual tokens | 0.4883 | 0.5352 |
| Zero visual tokens | 0.0020 | 0.0000 |
| Random visual tokens | 0.0020 | 0.0000 |
| Shuffled visual tokens | 0.0117 | 0.0215 |
These controls show that the learned visual tokens materially affect the image representation.
Matching
Cyclic-negative matching AUROC is 0.9733 on MONET. Text-nearest hard-negative AUROC
is 0.4961, so the current model is substantially better at coarse compatibility than
fine-grained semantic discrimination.
Compression details and size comparison
The released LFM2.5 body uses GPTQ INT4 W4A16 compression:
| Compression setting | Value |
|---|---|
| Method | GPTQ INT4 |
| Weight/activation format | W4A16 |
| Quantized layers | Linear layers; lm_head excluded |
| Block size | 128 |
| Activation order | Static |
| Dampening fraction | 0.01 |
| Calibration | 256 fused multimodal inputs_embeds samples |
| Maximum calibration sequence length | 160 tokens |
| Compressor revision | 8cec0acc1931de6f8f73257151ab7007c14dbf4e |
The calibration samples contained actual projected visual tokens rather than text-only inputs, so the quantization pass reflects the multimodal path.
Storage comparison
| Artifact | Model weights | Full package | Relative to original package |
|---|---|---|---|
| Original LFM2.5 Encoder 230M FP32 | 918.79 MB | 923.65 MB | 100% |
| Clean BF16 reference | 459.40 MB | 474.81 MB | 51.4% |
| Clean GPTQ INT4 release | 355.03 MB | 370.46 MB | 40.1% |
The GPTQ release is approximately 59.9% smaller than the original package and 22.0% smaller than the clean BF16 package. The external SigLIP2 vision tower is not included in these sizes and is downloaded separately.
Fidelity and runtime
| Metric | GPTQ INT4 |
|---|---|
| Hidden cosine versus BF16 | 0.9427 |
| Logit cosine versus BF16 | 0.9654 |
| Multimodal latency, batch 4 | 45.7 ms on NVIDIA L4 |
| Peak multimodal VRAM | 2.38 GiB on NVIDIA L4 |
The current GPTQ backend is intended primarily for weight-size reduction. Native BF16 is faster and uses less runtime VRAM in the tested environment.
Limitations
- Encoder-only: no generation is supported.
- Hard-negative matching remains near chance.
- Masked-token prediction does not currently improve over text-only or zero-image controls.
- External evaluation is a bounded CC3M subset rather than a full benchmark.
- The SigLIP2 vision tower remains an external Hugging Face dependency.
Reproducibility
The evaluation/ directory contains the saved BF16, GPTQ, and external CC3M metrics,
retrieval summary, matching results, ablations, robustness results, and the detailed
deep-evaluation summary.
The compact release corresponds to the clean image-only training branch and should be used instead of the earlier paired-caption artifact.
- Downloads last month
- 27