| --- |
| license: other |
| license_name: lfm1.0 |
| license_link: https://huggingface.co/LiquidAI/LFM2.5-230M/blob/main/LICENSE |
| base_model: LiquidAI/LFM2.5-230M |
| tags: |
| - qualcomm |
| - hexagon |
| - npu |
| - qnn |
| - qhexrt |
| - on-device |
| - lfm2 |
| language: |
| - en |
| pipeline_tag: text-generation |
| --- |
| |
| # LFM2.5-230M — Hexagon NPU (QHexRT) bundle |
|
|
| [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M) compiled to run on the **Qualcomm Hexagon |
| v79 NPU** (Snapdragon 8 Elite / SM8750, e.g. Galaxy S25) via the **[QHexRT](https://github.com/RunanywhereAI/QHexRT)** |
| runtime. Pure on-device inference — **no Python in the hot path.** |
|
|
| 14-layer hybrid model (6 GQA-attention + 8 short-conv), hidden 1024, vocab 65536, tied lm-head. Runs **W8 |
| weight-only** (int8 weights, fp16 activations) with **GQA-native decode** + **batched prefill** + an |
| **on-NPU lm-head**. Greedy output matches the HF model exactly (`"The capital of France is"` → `" Paris."`). |
|
|
| ## Measured (Samsung S25, Hexagon v79) |
|
|
| | config | decode | prefill | peak RAM | |
| |---|---|---|---| |
| | **MAXCTX 512** (`lfm2-5-230m.json`) | **164 tok/s** (6.1 ms/tok) | batched, ~17k tok/s (≤512 prompt) | ~430 MB | |
| | **MAXCTX 2048** (`lfm2-5-230m-2048.json`) | **127 tok/s** (7.9 ms/tok) | **~8,950 tok/s** (2k-token prompt in 227 ms) | ~440 MB | |
|
|
| For reference, LiquidAI's published **S25 CPU (int4)** numbers on 2k input are **prefill 1158 / decode 213 |
| tok/s** — this NPU bundle does prefill **~7.7× faster**, at far lower power. (Decode is W8 here: 4-bit weights |
| are blocked by the v79 HTP toolchain, so W8 is the floor; the NPU still wins prefill and frees the CPU.) |
|
|
| ## Contents (`v79/`) |
|
|
| | file | what | |
| |---|---| |
| | `lfm230_dec_512_w8.bin` / `lfm230_dec_2048_w8.bin` | W8 GQA-native decode (MAXCTX 512 / 2048) | |
| | `lfm230_pf_512_w8.bin` / `lfm230_pf_2048_w8.bin` | W8 batched prefill (PN 512 / 2048) | |
| | `lfm230_lmh_w8.bin` | W8 tied lm-head `hidden[1,1024]→logits[1,65536]` on-NPU | |
| | `lfm_embed_f16.bin` | tied embedding table (host token→hidden lookup) | |
| | `tokenizer.json` | the LFM2.5 tokenizer | |
| | `lfm2-5-230m.json` / `lfm2-5-230m-2048.json` | QHexRT manifests (512 / 2048; declare the 14-layer `attn_idx`/`conv_idx` schedule) | |
|
|
| ## Run |
|
|
| ```bash |
| hf download runanywhere/lfm2_5_230m_HNPU --local-dir lfm2_5_230m_HNPU |
| adb push lfm2_5_230m_HNPU/v79 /data/local/tmp/lfm230 # PowerShell + native paths on Windows |
| adb shell "cd /data/local/tmp/lfm230 && LD_LIBRARY_PATH=. \ |
| ./qhx_generate lfm2-5-230m-2048.json libQnnHtp.so libQnnSystem.so . 64 'The capital of France is'" |
| ``` |
|
|
| (One-time: stage the QAIRT v79 runtime libs + the `qhx_generate` tool into the same dir — see the QHexRT |
| deploy docs.) Arch-pinned to **v79**; a v79 binary will not load on other Hexagon arches. |
|
|