Qwen3.5-2B-PARO / README.md
liang2kl's picture
Upload README.md with huggingface_hub
2f33363 verified
|
Raw
History Blame
3.17 kB
---
library_name: transformers
license: apache-2.0
pipeline_tag: image-text-to-text
base_model:
- Qwen/Qwen3.5-2B
---
# z-lab/Qwen3.5-2B-PARO
**Pairwise Rotation Quantization for Efficient Reasoning LLM Inference**
<p>
<a href="https://arxiv.org/abs/2511.10645"><img src="https://img.shields.io/badge/arXiv-2511.10645-b31b1b.svg" alt="Paper"></a>
<a href="https://paroquant.z-lab.ai"><img src="https://img.shields.io/badge/Blog-ParoQuant-blue" alt="Blog"></a>
<a href="https://huggingface.co/collections/z-lab/paroquant"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Models-yellow" alt="Models"></a>
<a href="https://pypi.org/project/paroquant/"><img src="https://img.shields.io/pypi/v/paroquant" alt="PyPI"></a>
</p>
ParoQuant is the state-of-the-art INT4 quantization for LLMs. It closes the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX). For more information, see https://github.com/z-lab/paroquant.
z-lab/Qwen3.5-2B-PARO is a 4-bit [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) quantized with ParoQuant. Check out other ParoQuant models from the Hugging Face [collection](https://huggingface.co/collections/z-lab/paroquant).
## Quick Start
### Installation
```bash
# NVIDIA GPU (CUDA 12.9)
pip install "paroquant[vllm]"
# NVIDIA GPU (CUDA 13.0)
pip install "paroquant[vllm]" "vllm==0.19.0" \
--extra-index-url https://wheels.vllm.ai/2a69949bdadf0e8942b7a1619b229cb475beef20/cu130 \
--extra-index-url https://download.pytorch.org/whl/cu130
# Apple Silicon
pip install "paroquant[mlx]"
```
### Interactive Chat
```bash
python -m paroquant.cli.chat --model z-lab/Qwen3.5-2B-PARO
```
### OpenAI-Compatible API Server
```bash
python -m paroquant.cli.serve --model z-lab/Qwen3.5-2B-PARO --port 8000
```
For vLLM, the arguments are passed to the vLLM server directly. See [vLLM docs](https://docs.vllm.ai/en/latest/configuration/serve_args/) for more details.
For MLX, add `--vlm` if you wish to load the VLM components and use the model's multimodal features. For vLLM, VLM components are loaded by default and can be skipped with the server argument `--language-model-only`.
### Docker (NVIDIA GPU)
> [!NOTE]
> The following commands map the local cache directory to the container in order to persist kernel cache across runs. Remove `-v ...` to disable this behaviour.
```bash
# Interactive chat
docker run --pull=always --rm -it --gpus all --ipc=host \
-v $HOME/.cache/paroquant:/root/.cache/paroquant \
ghcr.io/z-lab/paroquant:chat --model z-lab/Qwen3.5-2B-PARO
# API server (port 8000)
docker run --pull=always --rm -it --gpus all --ipc=host -p 8000:8000 \
-v $HOME/.cache/paroquant:/root/.cache/paroquant \
ghcr.io/z-lab/paroquant:serve --model z-lab/Qwen3.5-2B-PARO
```
## Citation
```bibtex
@inproceedings{liang2026paroquant,
title = {{ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference}},
author = {Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}
```