Qwen3.5-0.8B

This version of Qwen3.5-0.8B has been converted to run on the Axera NPU using w8a16 quantization.

Compatible with Pulsar2 version: 5.0

Convert tools links:

For those who are interested in model conversion, you can try to export axmodel through the original repo :

Pulsar2 Link, How to Convert LLM from Huggingface to axmodel

AXera NPU HOST LLM Runtime

Support Platform

Image Process

Chips Input Size Image Num TTFT (168 tokens) Throughput (w8a16) CMM Memory Flash Memory
AX650 384×384 1 282 ms 18.5 tokens/sec 1.27 GiB 1.54 GiB

Video Process

Chips Input Size Image Num TTFT (600 tokens) Throughput (w8a16) CMM Memory Flash Memory
AX650 384×384 8 706 ms 18.5 tokens/sec 1.27 GiB 1.54 GiB

The DDR capacity refers to the CMM memory that needs to be consumed. Ensure that the CMM memory allocation on the development board is greater than this value.

How to use

安装 axllm

方式一:克隆仓库后执行安装脚本:

git clone -b axllm https://github.com/AXERA-TECH/ax-llm.git
cd ax-llm
./install.sh

方式二:一行命令安装(默认分支 axllm):

curl -fsSL https://raw.githubusercontent.com/AXERA-TECH/ax-llm/axllm/install.sh | bash

方式三:下载Github Actions CI 导出的可执行程序(适合没有编译环境的用户):

如果没有编译环境,请到: https://github.com/AXERA-TECH/ax-llm/actions?query=branch%3Aaxllm 下载 最新 CI 导出的可执行程序axllm),然后:

chmod +x axllm
sudo mv axllm /usr/bin/axllm

模型下载(Hugging Face)

先创建模型目录并进入,然后下载到该目录:

mkdir -p AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047
cd AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047
hf download AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 --local-dir .

# structure of the downloaded files
tree -L 3
`-- AXERA-TECH
    `-- Qwen3.5-0.8B-AX650-C128-P1152-CTX2047
        |-- qwen3_5_vision.axmodel
        |-- README.md
        |-- config.json
        |-- image.png
        |-- model.embed_tokens.weight.bfloat16.bin
        |-- post_config.json
        |-- qwen3_5_tokenizer.txt
        |-- qwen3_5_text_p128_l0_together.axmodel
        ...
        |-- qwen3_5_text_p128_l23_together.axmodel
        |-- qwen3_5_text_post.axmodel
        `-- vision_cache

3 directories, 39 files

Inference with AX650 Host, such as M4N-Dock(爱芯派Pro) or AX650N DEMO Board

运行(CLI)

root@ax650 ~/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047 # ./axllm run .
10:34:22.765 INF Init:218 | LLM init start
10:34:22.766 INF Init:226 | mixed attention enabled: full_attention_interval=4 ref_full_layer_idx=3
tokenizer_type = 3
 37% | ###########                      |  10 /  27 [6.61s<17.84s, 1.51 count/s] init 8 axmodel ok,remain_cmm(4623 MB)^C
 96% | ##############################   |  26 /  27 [18.42s<19.13s, 1.41 count/s] init post axmodel ok,remain_cmm(3780 MB)
10:34:41.183 INF Init:368 | max_token_len : 2047
10:34:41.183 INF Init:371 | kv_cache_size : 512, kv_cache_num: 2047
10:34:41.183 INF Init:374 | prefill_token_num : 128
10:34:41.183 INF Init:379 | grp: 1, prefill_max_kv_cache_num : 1
10:34:41.183 INF Init:379 | grp: 2, prefill_max_kv_cache_num : 128
10:34:41.183 INF Init:379 | grp: 3, prefill_max_kv_cache_num : 256
10:34:41.183 INF Init:379 | grp: 4, prefill_max_kv_cache_num : 384
10:34:41.183 INF Init:379 | grp: 5, prefill_max_kv_cache_num : 512
10:34:41.184 INF Init:379 | grp: 6, prefill_max_kv_cache_num : 640
10:34:41.184 INF Init:379 | grp: 7, prefill_max_kv_cache_num : 768
10:34:41.184 INF Init:379 | grp: 8, prefill_max_kv_cache_num : 896
10:34:41.184 INF Init:379 | grp: 9, prefill_max_kv_cache_num : 1024
10:34:41.184 INF Init:379 | grp: 10, prefill_max_kv_cache_num : 1152
10:34:41.184 INF Init:384 | prefill_max_token_num : 1152
10:34:41.184 INF Init:27 | LLaMaEmbedSelector use mmap
100% | ################################ |  27 /  27 [18.42s<18.42s, 1.47 count/s] embed_selector init ok
10:34:42.604 INF Init:643 | Qwen-VL token ids: vision_start=248053 image_pad=248056 video_pad=248057
10:34:42.604 INF Init:668 | VisionModule init ok: type=Qwen3VL, tokens_per_block=144, embed_size=1024, out_dtype=fp32
10:34:42.604 WRN Init:677 | Vision preprocess backend: SimpleCV (OpenCV not found at build time; minor differences vs OpenCV are possible)
10:34:42.609 INF load_config:282 | load config: 
10:34:42.609 INF load_config:282 | {
10:34:42.609 INF load_config:282 |     "enable_repetition_penalty": false,
10:34:42.609 INF load_config:282 |     "enable_temperature": false,
10:34:42.609 INF load_config:282 |     "enable_top_k_sampling": true,
10:34:42.609 INF load_config:282 |     "enable_top_p_sampling": false,
10:34:42.609 INF load_config:282 |     "penalty_window": 20,
10:34:42.609 INF load_config:282 |     "repetition_penalty": 1.2,
10:34:42.609 INF load_config:282 |     "temperature": 0.9,
10:34:42.609 INF load_config:282 |     "top_k": 10,
10:34:42.609 INF load_config:282 |     "top_p": 0.8
10:34:42.609 INF load_config:282 | }
10:34:42.609 INF Init:448 | LLM init ok
Commands:
  /q, /exit  退出
  /reset     重置 kvcache
  /dd        删除一轮对话
  /pp        打印历史对话
Ctrl+C: 停止当前生成
VLM enabled: after each prompt, input image path (empty = text-only). Use "video:<frames_dir>" for video.
----------------------------------------
prompt >> describe the image
image >> image.png
10:35:26.924 INF EncodeForContent:973 | Qwen-VL pixel_values[0] bytes=884736 min=0 max=238 (w=384 h=384 tp=2 ps=16 sm=2)
10:35:26.970 INF EncodeForContent:996 | vision cache store: image.png
10:35:27.004 INF SetKVCache:747 | prefill_grpid:3 kv_cache_num:256 precompute_len:0 input_num_token:168
10:35:27.004 INF SetKVCache:749 | current prefill_max_token_num:1152
10:35:27.004 INF SetKVCache:750 | first run
10:35:27.046 INF Run:805 | input token num : 168, prefill_split_num : 2
10:35:27.046 INF Run:845 | prefill chunk p=0 history_len=0 grpid=1 kv_cache_num=0 input_tokens=128
10:35:27.046 INF Run:868 | prefill indices shape: p=0 idx_elems=128 idx_rows=1 pos_rows=3
10:35:27.178 INF Run:845 | prefill chunk p=1 history_len=128 grpid=2 kv_cache_num=128 input_tokens=40
10:35:27.178 INF Run:868 | prefill indices shape: p=1 idx_elems=128 idx_rows=1 pos_rows=3
10:35:27.327 INF Run:1010 | ttft: 281.58 ms
<think>

</think>

This is a surreal, digitally rendered image that blends elements of sci-fi and fantasy.

**Setting & Atmosphere:**
The scene takes place in an alien, forest-like environment. Tall, misty trees rise in the background, and the ground is covered in lush, pale green and gray foliage, suggesting a dense jungle or canopy. The color palette is desaturated, using greys, whites, and muted greens, which enhances the otherworldly and sterile feel.

**Subject:**
The central figure is an astronaut in a spacesuit and a helmet, standing in the foreground. The astronaut is positioned on two legs (a human pose) rather than on a standing base, creating an uncanny or surreal effect. Their posture is alert and upright, as if they are actively observing or preparing for action within their new environment.

**Composition & Style:**
*   **Composition:** The astronaut is placed centrally but appears to be emerging from or integrated into the dense background, which can be described as a thick layer of foliage or mist. The framing creates a sense of depth and enclosure.
*   **Style:** The image has a painterly or textured quality, with visible brushstrokes or grain that give it a handmade feel, contrasting with the clean, technical render of the astronaut's helmet.

**Overall Impression:**
The image creates a sense of wonder or disorientation—the astronaut, with their alien habitat, stands out as both a strange and beautiful figure. It evokes a mood of quiet observation in a vast, otherworldly wilderness.

10:35:44.617 NTC Run:1132 | hit eos,avg 18.51 token/s
10:35:44.621 INF GetKVCache:721 | precompute_len:356, remaining:796

启动服务(OpenAI 兼容)

root@ax650:~# axllm serve AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047
[I][                            Init][ 138]: LLM init start
tokenizer_type = 1
 96% | ███████████████████████████████   |  30 /  31 [4.63s<4.79s, 6.47 count/s] init post axmodel ok,remain_cmm(9563 MB)
[I][                            Init][ 199]: max_token_len : 2047
[I][                            Init][ 202]: kv_cache_size : 1024, kv_cache_num: 2047
[I][                            Init][ 205]: prefill_token_num : 128
[I][                            Init][ 209]: grp: 1, prefill_max_kv_cache_num : 1
[I][                            Init][ 209]: grp: 2, prefill_max_kv_cache_num : 128
[I][                            Init][ 209]: grp: 3, prefill_max_kv_cache_num : 256
[I][                            Init][ 209]: grp: 4, prefill_max_kv_cache_num : 384
[I][                            Init][ 209]: grp: 5, prefill_max_kv_cache_num : 512
[I][                            Init][ 209]: grp: 6, prefill_max_kv_cache_num : 640
[I][                            Init][ 209]: grp: 7, prefill_max_kv_cache_num : 768
[I][                            Init][ 209]: grp: 8, prefill_max_kv_cache_num : 896
[I][                            Init][ 209]: grp: 9, prefill_max_kv_cache_num : 1024
[I][                            Init][ 209]: grp: 10, prefill_max_kv_cache_num : 1152
[I][                            Init][ 214]: prefill_max_token_num : 1152
[I][                            Init][  27]: LLaMaEmbedSelector use mmap
100% | ████████████████████████████████ |  31 /  31 [4.64s<4.64s, 6.69 count/s] embed_selector init ok
[W][                            Init][ 457]: Qwen-VL vision size override: cfg=448x448 bytes=1204224, model_input_bytes=884736 -> 384x384 (square).
[I][                            Init][ 641]: Qwen-VL token ids: vision_start=151652 image_pad=151655 video_pad=151656
[I][                            Init][ 666]: VisionModule init ok: type=Qwen3VL, tokens_per_block=144, embed_size=2048, out_dtype=fp32
[I][                            Init][ 672]: VisionModule deepstack enabled: layers=3
[I][                     load_config][ 282]: load config:
{
    "enable_repetition_penalty": false,
    "enable_temperature": false,
    "enable_top_k_sampling": false,
    "enable_top_p_sampling": false,
    "penalty_window": 20,
    "repetition_penalty": 1.2,
    "temperature": 0.9,
    "top_k": 10,
    "top_p": 0.8
}

[I][                            Init][ 272]: LLM init ok
Starting server on port 8000 with model 'AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047'...
OpenAI API Server starting on http://0.0.0.0:8000
Max concurrency: 1
Models: AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047

OpenAI 调用示例

from openai import OpenAI

API_URL = "http://127.0.0.1:8000/v1"
MODEL = "AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047"

messages = [
    {"role": "system", "content": [{"type": "text", "text": "you are a helpful assistant."}]},
    {"role": "user", "content": "hello"},
]

client = OpenAI(api_key="not-needed", base_url=API_URL)
completion = client.chat.completions.create(
    model=MODEL,
    messages=messages,
)

print(completion.choices[0].message.content)

OpenAI 流式调用示例

from openai import OpenAI

API_URL = "http://127.0.0.1:8000/v1"
MODEL = "AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047"

messages = [
    {"role": "system", "content": [{"type": "text", "text": "you are a helpful assistant."}]},
    {"role": "user", "content": "hello"},
]

client = OpenAI(api_key="not-needed", base_url=API_URL)
stream = client.chat.completions.create(
    model=MODEL,
    messages=messages,
    stream=True,
)

print("assistant:")
for ev in stream:
    delta = getattr(ev.choices[0], "delta", None)
    if delta and getattr(delta, "content", None):
        print(delta.content, end="", flush=True)
print("
")
Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047

Finetuned
(276)
this model

Collection including AXERA-TECH/Qwen3.5-0.8B-AX650-C128-P1152-CTX2047