license: apache-2.0
tags:
- qhexrt
- internvl
- vlm
- image-text-to-text
- hexagon
- npu
- qnn
base_model: OpenGVLab/InternVL3_5-1B
pipeline_tag: image-text-to-text
language:
- en
InternVL3.5-1B — QHexRT NPU bundle
Precompiled InternVL3.5-1B vision-language model for the QHexRT runtime on Qualcomm Hexagon NPU.
Image → caption runs entirely on the NPU: the InternViT-300M vision encoder and the full Qwen3-0.6B
language model are qnn_context_binary graphs; the host does only light glue (image resize, patch-embed,
pixel-shuffle projector, tokenize/splice, sampling). Device-validated end-to-end on diverse photos.
The model is driven by QHexRT's declarative manifest plan-interpreter (Session::generate via the
internvl_generate host-op) — the same engine that runs the other *_HNPU bundles.
Context binaries are arch-pinned — pick the sibling dir that matches your device.
Architectures
| Dir | Device (example) | Status |
|---|---|---|
v79/ |
Galaxy S25 / SM8750 (Hexagon v79) | Device-validated |
v81/ |
SM8850 (Hexagon v81) | Device-validated |
v75/ |
Galaxy S24 / SM8650 (Hexagon v75) | Device-validated |
Run (QHexRT CLI)
hf download runanywhere/internvl3_5_1b_HNPU --local-dir internvl3_5_1b_HNPU
# qhx_generate from a QHexRT build; QNN libs from QAIRT SDK (lib/aarch64-android) + matching HTP skel.
adb push internvl3_5_1b_HNPU/v79 /data/local/tmp/wq/internvl # or v75/ or v81/
adb push my_photo.jpg /data/local/tmp/wq/internvl/
adb shell "cd /data/local/tmp/wq && export ADSP_LIBRARY_PATH='/data/local/tmp/wq/dsp;/data/local/tmp/wq;/vendor/dsp/cdsp'; \
LD_LIBRARY_PATH=. ./qhx_generate internvl/internvl3_5-1b-512.json libQnnHtp.so libQnnSystem.so internvl \
60 'What animals are in this picture and what are they doing?' internvl/my_photo.jpg"
# -> text: "In the image, there are two cats lying on a pink couch ..."
- Image: any aspect ratio / format (JPEG/PNG/…); center-cropped + resized to 448×448 host-side (256 image tokens / 1 tile).
- Arbitrary prompts: the padded
AR=512prefill accepts any prompt up to ~256 text tokens (image is fixed at 256 tokens). Ask real questions, not just a fixed caption string.
Contents (v79/)
| file | what | size |
|---|---|---|
internvl3_5-1b-512.json |
QHexRT manifest (vlm family, internvl_generate plan) |
— |
ivl_enc2.bin |
InternViT-300M vision encoder — 24-layer, fp16, 2D [seq,dim] I/O |
621 MB |
ivl_pf_512.bin |
Qwen3-0.6B batched prefill graph, AR=512 (padded) — arbitrary prompts up to ~256 text tokens (fp16) | 890 MB |
ivl_dec_w4.bin |
Qwen3-0.6B decode step (1 token), W4A16 | 225 MB |
ivl_lmh_q.bin |
untied lm-head, W8A16 (int8) → logits over 151936 vocab | 158 MB |
ivl_embed_f16.bin |
token embedding table [151936,1024] f16 (mmap, host lookup) |
311 MB |
tokenizer.json |
Qwen3 tokenizer (vocab 151936) | 11 MB |
vlm/ |
host vision preproc: Conv2d patch-embed, pos-embed, layer-norm + pixel-shuffle MLP projector, prompt template | 27 MB |
Total ≈ 2.25 GB on disk; ≈ 2.0 GB peak device RSS at runtime (weights are ion-shared host↔DSP, not duplicated).
Performance (measured, v79 / S25)
| stage | time |
|---|---|
| vision encode (one-time) | 181 ms |
| prefill / TTFT | 101 ms |
| decode | ~28 ms/token on the NPU graphs |
End-to-end wall time includes single-threaded host preproc (patch-embed + projector, ~1.5 s one-time); the NPU compute itself is the table above.
v75 (SM8650 / soc_model 57)
Device-validated on Galaxy S24 (Hexagon v75): all four context binaries load; image→caption runs end-to-end. Vision 206 ms, prefill TTFT 139 ms, decode ~23 ms/tok (QAIRT 2.47, soc 57).
Same graph layout as v79/ (ivl_enc2, ivl_pf_512, ivl_dec_w4, ivl_lmh_q). Vision + prefill are
fp16; lm-head is W8 weight-only. Decode graph name is ivl_dec_w4 but weights are W8
weight-only (per-row int8 via dump_wqo_params, not true W4 like v79/) — same compile workaround used
for other v75 LLM bundles on Apple Silicon hosts where qairt-quantizer SIGILLs under Docker.
Files (v75/)
| file | what | size |
|---|---|---|
internvl3_5-1b-512.json |
QHexRT manifest (dsp_arch: v75) |
— |
ivl_enc2.bin |
InternViT-300M vision encoder (fp16) | 592 MB |
ivl_pf_512.bin |
Batched LLM prefill Pn=512 (fp16) | 855 MB |
ivl_dec_w4.bin |
GQA-native decode MAXCTX=512 (W8 weight-only, graph name ivl_dec_w4) |
424 MB |
ivl_lmh_q.bin |
Untied lm-head (W8 weight-only) | 151 MB |
ivl_embed_f16.bin |
Token embedding table f16 | 297 MB |
ivl_tokenizer.json |
Qwen2 tokenizer — separate filename so it does not collide with other QHexRT bundles on device | 11 MB |
vlm/ |
Host patch-embed + pixel-shuffle projector + prompt template | 27 MB |
Total ≈ 2.3 GB.
adb push internvl3_5_1b_HNPU/v75 /data/local/tmp/wq/internvl
adb push my_photo.jpg /data/local/tmp/wq/internvl/
adb shell "cd /data/local/tmp/wq && export ADSP_LIBRARY_PATH='/data/local/tmp/wq/dsp;/data/local/tmp/wq;/vendor/dsp/cdsp'; \
LD_LIBRARY_PATH=. ./qhx_generate internvl/internvl3_5-1b-512.json libQnnHtp.so libQnnSystem.so internvl \
60 'Describe this image in detail.' internvl/my_photo.jpg"
v81 (SM8850 / soc_model 87)
Device-validated on SM8850: image→caption tracks the image — e.g. an orange dahlia ->
"The image features a close-up of a single, large, orange flower with a textured, layered appearance. The
flower has a circular shape with numerous petals radiating from a central point...". Vision 147 ms, prefill
TTFT 115 ms, decode 18.4 ms/tok. The v81 recipe was completed by reconstructing the two graphs the
internvl_generate host-op needs but the recipe lacked: a batched-prefill graph (ivl_pf, AR=512, GQA) and a
lm-head graph (ivl_lmh) — both gated vs the HF gold first token. All graphs f16 (f16 I/O), DLC route + O3.
Files (v81/)
internvl3_5-1b.json · ivl_enc_f16.bin (InternViT-300M vision) · ivl_pf_f16.bin (batched prefill) ·
ivl_dec_f16.bin (AR=1 decode) · ivl_lmh_f16.bin (lm-head) · ivl_embed_f16.bin · tokenizer.json ·
vlm/ (host patch/pos-embed + pixel-shuffle MLP projector weights + prompt template).
Notes
- Context binaries are dsp-arch-pinned — use the sibling
<arch>/dir that matches your device. - No custom op-package needed — all graphs are pure-native HTP (HVX/HMX).
- Precision (v79): vision fp16, prefill fp16, decode W4A16, lm-head W8A16.
- HTP gotcha baked in: all graph I/O is 2D
[seq,dim]— HTP mis-lays-out 3D[1,seq,dim]tensors. The Conv2d patch-embed and pixel-shuffle are host-side for the same reason. - Source model:
OpenGVLab/InternVL3_5-1B(InternViT-300M + Qwen3-0.6B).