kucukkanat's picture
Update model card
be5119b verified
|
Raw
History Blame Contribute Delete
3.3 kB
---
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M/blob/main/LICENSE
base_model: LiquidAI/LFM2.5-Encoder-350M
base_model_relation: quantized
library_name: transformers.js
pipeline_tag: fill-mask
tags:
- onnx
- transformers.js
- lfm2
- quantized
language:
- en
- de
- es
- fr
- it
- nl
- pl
- pt
- ar
- hi
- ja
- ru
- tr
- vi
- zh
---
# LFM2.5-Encoder-350M-ONNX
ONNX export of [`LiquidAI/LFM2.5-Encoder-350M`](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M), quantized to run **fully in the browser** through
[transformers.js](https://github.com/huggingface/transformers.js). No inference server: the weights are
fetched once, cached, and every forward pass happens in the tab.
The base bidirectional encoder with its tied masked-LM head. Outputs vocabulary logits and the
final hidden states, so it doubles as a sentence-embedding backbone.
All credit for the model itself goes to [Liquid AI](https://huggingface.co/LiquidAI). This repository
contains only a re-export; the weights are unchanged apart from quantization, and the original
[LFM Open License v1.0](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M/blob/main/LICENSE) applies.
**[Try it in your browser →](https://kucukkanat.github.io/lfm-encoders/)** — no install, no API key.
![Fill-mask running in the browser](https://raw.githubusercontent.com/kucukkanat/lfm-encoders/main/docs/screenshots/fill-mask.png)
Tooling, demo and the export pipeline: <https://github.com/kucukkanat/lfm-encoders>
## Files
| dtype | File | Size |
| --- | --- | --: |
| `q8` | `onnx/model_quantized.onnx` | 424 MB |
| `q4` | `onnx/model_q4.onnx` | 491 MB |
The graph takes `input_ids` + `attention_mask`, is dynamic in batch and sequence, and returns
`logits` and `last_hidden_state`.
## Usage
```js
import { AutoTokenizer, PreTrainedModel, Tensor } from "@huggingface/transformers";
const id = "kucukkanat/LFM2.5-Encoder-350M-ONNX";
const tokenizer = await AutoTokenizer.from_pretrained(id);
const model = await PreTrainedModel.from_pretrained(id, { dtype: "q8" });
const { input_ids } = tokenizer("some text");
const out = await model({
input_ids,
attention_mask: new Tensor("int64", new BigInt64Array(input_ids.dims[1]).fill(1n), input_ids.dims),
});
```
`PreTrainedModel` rather than `AutoModel` is deliberate: this is a plain "feed the named inputs, read the
named outputs" session, not one of transformers.js's built-in architectures.
## Accuracy
Measured from JavaScript against the fp32 PyTorch reference. Δ is the largest absolute difference in a
final probability.
| dtype | max Δ | mean Δ | top-5 disagreements (3 cases) |
| --- | --: | --: | --: |
| `fp32` | 9.6e-5 | 4.1e-5 | 0 |
| `q8` | 0.1846 | 0.1235 | 4 |
| `q4` | 0.2180 | 0.0978 | 3 |
## Notes
- `q8` is smaller on disk but uses **more** browser RAM than fp32 and runs slower: onnxruntime's WASM
kernels compute in float, so quantized weights are unpacked at session load. Quantization here buys
download size, not speed or memory.
- Budget roughly 1.5 GB of RAM per resident model, and expect a tab to hold its high-water mark until
reloaded.
- `fp16` / `q4f16` are deliberately absent: RMSNorm's variance overflows fp16 on this architecture and
every hidden state collapses to zeros.