--- license: mit language: - en - zh base_model: - Qwen/Qwen3.5-0.8B pipeline_tag: image-text-to-text library_name: transformers tags: - Qwen3.5 - Qwen3.5-0.8B - VLM --- # Qwen3.5-0.8B This version of Qwen3.5-0.8B has been converted to run on the Axera NPU using **w8a16** quantization. Compatible with Pulsar2 version: 5.0 ## Convert tools links: For those who are interested in model conversion, you can try to export axmodel through the original repo : - https://huggingface.co/Qwen/Qwen3.5-0.8B [Pulsar2 Link, How to Convert LLM from Huggingface to axmodel](https://pulsar2-docs.readthedocs.io/en/latest/appendix/build_llm.html) [AXera NPU HOST LLM Runtime](https://github.com/AXERA-TECH/ax-llm.git) ## Support Platform - AX650 - AX650N DEMO Board - [M4N-Dock(爱芯派Pro)](https://wiki.sipeed.com/hardware/zh/maixIV/m4ndock/m4ndock.html) - [M.2 Accelerator card](https://docs.m5stack.com/zh_CN/ai_hardware/LLM-8850_Card) **Image Process** | Chips | Input Size | Image Num | TTFT (168 tokens) | Throughput (w8a16) | CMM Memory | Flash Memory | |:------|:----------:|:---------:|------------------:|-------------------:|-----------:|-------------:| | AX650 | 384×384 | 1 | 282 ms | 18.5 tokens/sec | 1.27 GiB | 1.54 GiB | **Video Process** | Chips | Input Size | Image Num | TTFT (600 tokens) | Throughput (w8a16) | CMM Memory | Flash Memory | |:------|:----------:|:---------:|------------------:|-------------------:|-----------:|-------------:| | AX650 | 384×384 | 8 | 706 ms | 18.5 tokens/sec | 1.27 GiB | 1.54 GiB | The DDR capacity refers to the CMM memory that needs to be consumed. Ensure that the CMM memory allocation on the development board is greater than this value. ## How to use ## 安装 axllm 方式一:克隆仓库后执行安装脚本: ```shell git clone -b axllm https://github.com/AXERA-TECH/ax-llm.git cd ax-llm ./install.sh ``` 方式二:一行命令安装(默认分支 `axllm`): ```shell curl -fsSL https://raw.githubusercontent.com/AXERA-TECH/ax-llm/axllm/install.sh | bash ``` 方式三:下载Github Actions CI 导出的可执行程序(适合没有编译环境的用户): 如果没有编译环境,请到: `https://github.com/AXERA-TECH/ax-llm/actions?query=branch%3Aaxllm` 下载 **最新 CI 导出的可执行程序**(`axllm`),然后: ```shell chmod +x axllm sudo mv axllm /usr/bin/axllm ``` ## 模型下载(Hugging Face) 先创建模型目录并进入,然后下载到该目录: ```shell mkdir -p AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 cd AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 hf download AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 --local-dir . # structure of the downloaded files tree -L 3 `-- AXERA-TECH `-- Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 |-- qwen3_5_vision.axmodel |-- README.md |-- config.json |-- image.png |-- model.embed_tokens.weight.bfloat16.bin |-- post_config.json |-- qwen3_5_tokenizer.txt |-- qwen3_5_text_p128_l0_together.axmodel ... |-- qwen3_5_text_p128_l23_together.axmodel |-- qwen3_5_text_post.axmodel `-- vision_cache 3 directories, 39 files ``` ## Inference with AX650 Host, such as M4N-Dock(爱芯派Pro) or AX650N DEMO Board ### 运行(CLI) ```shell root@ax650 ~/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 # ./axllm run . 10:34:22.765 INF Init:218 | LLM init start 10:34:22.766 INF Init:226 | mixed attention enabled: full_attention_interval=4 ref_full_layer_idx=3 tokenizer_type = 3 37% | ########### | 10 / 27 [6.61s<17.84s, 1.51 count/s] init 8 axmodel ok,remain_cmm(4623 MB)^C 96% | ############################## | 26 / 27 [18.42s<19.13s, 1.41 count/s] init post axmodel ok,remain_cmm(3780 MB) 10:34:41.183 INF Init:368 | max_token_len : 2047 10:34:41.183 INF Init:371 | kv_cache_size : 512, kv_cache_num: 2047 10:34:41.183 INF Init:374 | prefill_token_num : 128 10:34:41.183 INF Init:379 | grp: 1, prefill_max_kv_cache_num : 1 10:34:41.183 INF Init:379 | grp: 2, prefill_max_kv_cache_num : 128 10:34:41.183 INF Init:379 | grp: 3, prefill_max_kv_cache_num : 256 10:34:41.183 INF Init:379 | grp: 4, prefill_max_kv_cache_num : 384 10:34:41.183 INF Init:379 | grp: 5, prefill_max_kv_cache_num : 512 10:34:41.184 INF Init:379 | grp: 6, prefill_max_kv_cache_num : 640 10:34:41.184 INF Init:379 | grp: 7, prefill_max_kv_cache_num : 768 10:34:41.184 INF Init:379 | grp: 8, prefill_max_kv_cache_num : 896 10:34:41.184 INF Init:379 | grp: 9, prefill_max_kv_cache_num : 1024 10:34:41.184 INF Init:379 | grp: 10, prefill_max_kv_cache_num : 1152 10:34:41.184 INF Init:384 | prefill_max_token_num : 1152 10:34:41.184 INF Init:27 | LLaMaEmbedSelector use mmap 100% | ################################ | 27 / 27 [18.42s<18.42s, 1.47 count/s] embed_selector init ok 10:34:42.604 INF Init:643 | Qwen-VL token ids: vision_start=248053 image_pad=248056 video_pad=248057 10:34:42.604 INF Init:668 | VisionModule init ok: type=Qwen3VL, tokens_per_block=144, embed_size=1024, out_dtype=fp32 10:34:42.604 WRN Init:677 | Vision preprocess backend: SimpleCV (OpenCV not found at build time; minor differences vs OpenCV are possible) 10:34:42.609 INF load_config:282 | load config: 10:34:42.609 INF load_config:282 | { 10:34:42.609 INF load_config:282 | "enable_repetition_penalty": false, 10:34:42.609 INF load_config:282 | "enable_temperature": false, 10:34:42.609 INF load_config:282 | "enable_top_k_sampling": true, 10:34:42.609 INF load_config:282 | "enable_top_p_sampling": false, 10:34:42.609 INF load_config:282 | "penalty_window": 20, 10:34:42.609 INF load_config:282 | "repetition_penalty": 1.2, 10:34:42.609 INF load_config:282 | "temperature": 0.9, 10:34:42.609 INF load_config:282 | "top_k": 10, 10:34:42.609 INF load_config:282 | "top_p": 0.8 10:34:42.609 INF load_config:282 | } 10:34:42.609 INF Init:448 | LLM init ok Commands: /q, /exit 退出 /reset 重置 kvcache /dd 删除一轮对话 /pp 打印历史对话 Ctrl+C: 停止当前生成 VLM enabled: after each prompt, input image path (empty = text-only). Use "video:" for video. ---------------------------------------- prompt >> describe the image image >> image.png 10:35:26.924 INF EncodeForContent:973 | Qwen-VL pixel_values[0] bytes=884736 min=0 max=238 (w=384 h=384 tp=2 ps=16 sm=2) 10:35:26.970 INF EncodeForContent:996 | vision cache store: image.png 10:35:27.004 INF SetKVCache:747 | prefill_grpid:3 kv_cache_num:256 precompute_len:0 input_num_token:168 10:35:27.004 INF SetKVCache:749 | current prefill_max_token_num:1152 10:35:27.004 INF SetKVCache:750 | first run 10:35:27.046 INF Run:805 | input token num : 168, prefill_split_num : 2 10:35:27.046 INF Run:845 | prefill chunk p=0 history_len=0 grpid=1 kv_cache_num=0 input_tokens=128 10:35:27.046 INF Run:868 | prefill indices shape: p=0 idx_elems=128 idx_rows=1 pos_rows=3 10:35:27.178 INF Run:845 | prefill chunk p=1 history_len=128 grpid=2 kv_cache_num=128 input_tokens=40 10:35:27.178 INF Run:868 | prefill indices shape: p=1 idx_elems=128 idx_rows=1 pos_rows=3 10:35:27.327 INF Run:1010 | ttft: 281.58 ms This is a surreal, digitally rendered image that blends elements of sci-fi and fantasy. **Setting & Atmosphere:** The scene takes place in an alien, forest-like environment. Tall, misty trees rise in the background, and the ground is covered in lush, pale green and gray foliage, suggesting a dense jungle or canopy. The color palette is desaturated, using greys, whites, and muted greens, which enhances the otherworldly and sterile feel. **Subject:** The central figure is an astronaut in a spacesuit and a helmet, standing in the foreground. The astronaut is positioned on two legs (a human pose) rather than on a standing base, creating an uncanny or surreal effect. Their posture is alert and upright, as if they are actively observing or preparing for action within their new environment. **Composition & Style:** * **Composition:** The astronaut is placed centrally but appears to be emerging from or integrated into the dense background, which can be described as a thick layer of foliage or mist. The framing creates a sense of depth and enclosure. * **Style:** The image has a painterly or textured quality, with visible brushstrokes or grain that give it a handmade feel, contrasting with the clean, technical render of the astronaut's helmet. **Overall Impression:** The image creates a sense of wonder or disorientation—the astronaut, with their alien habitat, stands out as both a strange and beautiful figure. It evokes a mood of quiet observation in a vast, otherworldly wilderness. 10:35:44.617 NTC Run:1132 | hit eos,avg 18.51 token/s 10:35:44.621 INF GetKVCache:721 | precompute_len:356, remaining:796 ``` ### 启动服务(OpenAI 兼容) ```shell root@ax650:~# axllm serve AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 [I][ Init][ 138]: LLM init start tokenizer_type = 1 96% | ███████████████████████████████ | 30 / 31 [4.63s<4.79s, 6.47 count/s] init post axmodel ok,remain_cmm(9563 MB) [I][ Init][ 199]: max_token_len : 2047 [I][ Init][ 202]: kv_cache_size : 1024, kv_cache_num: 2047 [I][ Init][ 205]: prefill_token_num : 128 [I][ Init][ 209]: grp: 1, prefill_max_kv_cache_num : 1 [I][ Init][ 209]: grp: 2, prefill_max_kv_cache_num : 128 [I][ Init][ 209]: grp: 3, prefill_max_kv_cache_num : 256 [I][ Init][ 209]: grp: 4, prefill_max_kv_cache_num : 384 [I][ Init][ 209]: grp: 5, prefill_max_kv_cache_num : 512 [I][ Init][ 209]: grp: 6, prefill_max_kv_cache_num : 640 [I][ Init][ 209]: grp: 7, prefill_max_kv_cache_num : 768 [I][ Init][ 209]: grp: 8, prefill_max_kv_cache_num : 896 [I][ Init][ 209]: grp: 9, prefill_max_kv_cache_num : 1024 [I][ Init][ 209]: grp: 10, prefill_max_kv_cache_num : 1152 [I][ Init][ 214]: prefill_max_token_num : 1152 [I][ Init][ 27]: LLaMaEmbedSelector use mmap 100% | ████████████████████████████████ | 31 / 31 [4.64s<4.64s, 6.69 count/s] embed_selector init ok [W][ Init][ 457]: Qwen-VL vision size override: cfg=448x448 bytes=1204224, model_input_bytes=884736 -> 384x384 (square). [I][ Init][ 641]: Qwen-VL token ids: vision_start=151652 image_pad=151655 video_pad=151656 [I][ Init][ 666]: VisionModule init ok: type=Qwen3VL, tokens_per_block=144, embed_size=2048, out_dtype=fp32 [I][ Init][ 672]: VisionModule deepstack enabled: layers=3 [I][ load_config][ 282]: load config: { "enable_repetition_penalty": false, "enable_temperature": false, "enable_top_k_sampling": false, "enable_top_p_sampling": false, "penalty_window": 20, "repetition_penalty": 1.2, "temperature": 0.9, "top_k": 10, "top_p": 0.8 } [I][ Init][ 272]: LLM init ok Starting server on port 8000 with model 'AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047'... OpenAI API Server starting on http://0.0.0.0:8000 Max concurrency: 1 Models: AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 ``` ### OpenAI 调用示例 ```python from openai import OpenAI API_URL = "http://127.0.0.1:8000/v1" MODEL = "AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047" messages = [ {"role": "system", "content": [{"type": "text", "text": "you are a helpful assistant."}]}, {"role": "user", "content": "hello"}, ] client = OpenAI(api_key="not-needed", base_url=API_URL) completion = client.chat.completions.create( model=MODEL, messages=messages, ) print(completion.choices[0].message.content) ``` ### OpenAI 流式调用示例 ```python from openai import OpenAI API_URL = "http://127.0.0.1:8000/v1" MODEL = "AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047" messages = [ {"role": "system", "content": [{"type": "text", "text": "you are a helpful assistant."}]}, {"role": "user", "content": "hello"}, ] client = OpenAI(api_key="not-needed", base_url=API_URL) stream = client.chat.completions.create( model=MODEL, messages=messages, stream=True, ) print("assistant:") for ev in stream: delta = getattr(ev.choices[0], "delta", None) if delta and getattr(delta, "content", None): print(delta.content, end="", flush=True) print(" ") ```