File size: 3,461 Bytes
552f786
 
 
 
2f33363
 
ee7357c
 
552f786
 
 
 
 
 
5f44948
552f786
 
 
 
 
 
9a5fe3b
552f786
c931afe
552f786
 
 
 
 
 
 
c931afe
552f786
 
c931afe
ee7357c
 
c931afe
 
552f786
 
 
 
 
 
 
 
 
 
 
 
58f3e1a
 
552f786
54a0883
552f786
 
58f3e1a
 
 
54a0883
58f3e1a
f854aa5
 
c931afe
58f3e1a
 
 
552f786
 
c931afe
54a0883
c931afe
552f786
 
 
c931afe
552f786
 
 
 
c931afe
552f786
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
---
library_name: transformers
license: apache-2.0
pipeline_tag: image-text-to-text
base_model:
- Qwen/Qwen3.5-2B
tags:
  - mlx
---

# z-lab/Qwen3.5-2B-PARO

**Pairwise Rotation Quantization for Efficient Reasoning LLM Inference**

<p>
  <a href="https://arxiv.org/abs/2511.10645"><img src="https://img.shields.io/badge/arXiv-2511.10645-b31b1b.svg" alt="Paper"></a>
  <a href="https://paroquant.z-lab.ai"><img src="https://img.shields.io/badge/Blog-ParoQuant-blue" alt="Blog"></a>
  <a href="https://huggingface.co/collections/z-lab/paroquant"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Models-yellow" alt="Models"></a>
  <a href="https://pypi.org/project/paroquant/"><img src="https://img.shields.io/pypi/v/paroquant" alt="PyPI"></a>
</p>

ParoQuant is the state-of-the-art INT4 quantization for LLMs. It closes the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX). For more information, see https://github.com/z-lab/paroquant.

z-lab/Qwen3.5-2B-PARO is a 4-bit [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) quantized with ParoQuant. Check out other ParoQuant models from the Hugging Face [collection](https://huggingface.co/collections/z-lab/paroquant).


## Quick Start

### Installation

```bash
# NVIDIA GPU (CUDA 12.9)
pip install "paroquant[vllm]"

# NVIDIA GPU (CUDA 13.0)
pip install "paroquant[vllm]" "vllm==0.19.1" \
  --extra-index-url https://wheels.vllm.ai/0.19.1/cu130 \
  --extra-index-url https://download.pytorch.org/whl/cu130

# Apple Silicon
pip install "paroquant[mlx]"
```

### Interactive Chat

```bash
python -m paroquant.cli.chat --model z-lab/Qwen3.5-2B-PARO
```

### OpenAI-Compatible API Server

For vLLM, you can directly use `vllm serve` to serve ParoQuant models:

```bash
vllm serve z-lab/Qwen3.5-2B-PARO --port 8000
```

For other frameworks:

```bash
python -m paroquant.cli.serve --model z-lab/Qwen3.5-2B-PARO --port 8000
```

For MLX, add `--vlm` if you wish to load the VLM components and use the model's multimodal features. For vLLM, VLM components are loaded by default and can be skipped with the server argument `--language-model-only`.

> [!NOTE]
> The visual components in this checkpoint is stored in original precision, and only the language components are quantized to 4 bits; as a result, the model size is larger than a fully-quantized model. Avoid loading the VLM components if you are not using the multimodal features for the best efficiency.

### Docker (NVIDIA GPU)

> [!NOTE]
> The following commands map the local cache directory to the container in order to persist kernel cache across runs. Remove `-v ...` to disable this behavior.

```bash
# Interactive chat
docker run --pull=always --rm -it --gpus all --ipc=host \
  -v $HOME/.cache/paroquant:/root/.cache/paroquant \
  ghcr.io/z-lab/paroquant:chat --model z-lab/Qwen3.5-2B-PARO

# API server (port 8000)
docker run --pull=always --rm -it --gpus all --ipc=host -p 8000:8000 \
  -v $HOME/.cache/paroquant:/root/.cache/paroquant \
  ghcr.io/z-lab/paroquant:serve --model z-lab/Qwen3.5-2B-PARO
```

## Citation

```bibtex
@inproceedings{liang2026paroquant,
  title     = {{ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference}},
  author    = {Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026}
}
```