--- license: other license_name: lfm1.0 license_link: https://huggingface.co/LiquidAI/LFM2.5-230M/blob/main/LICENSE base_model: LiquidAI/LFM2.5-230M tags: - qualcomm - hexagon - npu - qnn - qhexrt - on-device - lfm2 language: - en pipeline_tag: text-generation --- # LFM2.5-230M — Hexagon NPU (QHexRT) bundle [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M) compiled to run on the **Qualcomm Hexagon v79 NPU** (Snapdragon 8 Elite / SM8750, e.g. Galaxy S25) via the **[QHexRT](https://github.com/RunanywhereAI/QHexRT)** runtime. Pure on-device inference — **no Python in the hot path.** 14-layer hybrid model (6 GQA-attention + 8 short-conv), hidden 1024, vocab 65536, tied lm-head. Runs **W8 weight-only** (int8 weights, fp16 activations) with **GQA-native decode** + **batched prefill** + an **on-NPU lm-head**. Greedy output matches the HF model exactly (`"The capital of France is"` → `" Paris."`). ## Measured (Samsung S25, Hexagon v79) | config | decode | prefill | peak RAM | |---|---|---|---| | **MAXCTX 512** (`lfm2-5-230m.json`) | **164 tok/s** (6.1 ms/tok) | batched, ~17k tok/s (≤512 prompt) | ~430 MB | | **MAXCTX 2048** (`lfm2-5-230m-2048.json`) | **127 tok/s** (7.9 ms/tok) | **~8,950 tok/s** (2k-token prompt in 227 ms) | ~440 MB | For reference, LiquidAI's published **S25 CPU (int4)** numbers on 2k input are **prefill 1158 / decode 213 tok/s** — this NPU bundle does prefill **~7.7× faster**, at far lower power. (Decode is W8 here: 4-bit weights are blocked by the v79 HTP toolchain, so W8 is the floor; the NPU still wins prefill and frees the CPU.) ## Contents (`v79/`) | file | what | |---|---| | `lfm230_dec_512_w8.bin` / `lfm230_dec_2048_w8.bin` | W8 GQA-native decode (MAXCTX 512 / 2048) | | `lfm230_pf_512_w8.bin` / `lfm230_pf_2048_w8.bin` | W8 batched prefill (PN 512 / 2048) | | `lfm230_lmh_w8.bin` | W8 tied lm-head `hidden[1,1024]→logits[1,65536]` on-NPU | | `lfm_embed_f16.bin` | tied embedding table (host token→hidden lookup) | | `tokenizer.json` | the LFM2.5 tokenizer | | `lfm2-5-230m.json` / `lfm2-5-230m-2048.json` | QHexRT manifests (512 / 2048; declare the 14-layer `attn_idx`/`conv_idx` schedule) | ## Run ```bash hf download runanywhere/lfm2_5_230m_HNPU --local-dir lfm2_5_230m_HNPU adb push lfm2_5_230m_HNPU/v79 /data/local/tmp/lfm230 # PowerShell + native paths on Windows adb shell "cd /data/local/tmp/lfm230 && LD_LIBRARY_PATH=. \ ./qhx_generate lfm2-5-230m-2048.json libQnnHtp.so libQnnSystem.so . 64 'The capital of France is'" ``` (One-time: stage the QAIRT v79 runtime libs + the `qhx_generate` tool into the same dir — see the QHexRT deploy docs.) Arch-pinned to **v79**; a v79 binary will not load on other Hexagon arches.