Qwen3.5-2B-PARO / README.md
liang2kl's picture
Upload README.md with huggingface_hub
c931afe verified
|
Raw
History Blame
3.13 kB
metadata
library_name: transformers
license: apache-2.0
pipeline_tag: image-text-to-text
base_model:
  - Qwen/Qwen3.5-2B

z-lab/Qwen3.5-2B-PARO

Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

Paper Blog Models PyPI

ParoQuant is the state-of-the-art INT4 quantization for LLMs. It closes the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX).

z-lab/Qwen3.5-2B-PARO is a 4-bit Qwen/Qwen3.5-2B quantized with ParoQuant. Check out other ParoQuant models from the Hugging Face collection.

Quick Start

Installation

# NVIDIA GPU (CUDA 12.9)
pip install "paroquant[vllm]"

# NVIDIA GPU (CUDA 13.0)
pip install "paroquant[vllm] vllm==0.17.1" \
  --extra-index-url https://wheels.vllm.ai/0.17.1/cu130 \
  --extra-index-url https://download.pytorch.org/whl/cu130

# Apple Silicon
pip install "paroquant[mlx]"

Interactive Chat

python -m paroquant.cli.chat --model z-lab/Qwen3.5-2B-PARO

Add --llm-only if you do not wish to load the VLM components.

OpenAI-Compatible API Server

python -m paroquant.cli.serve --model z-lab/Qwen3.5-2B-PARO --port 8000

Add --llm-only if you do not wish to load the VLM components.

Agent with Tool Calling

Start the API server first, then install the agent dependencies and run:

pip install "paroquant[agent]"
python -m paroquant.cli.agent --model z-lab/Qwen3.5-2B-PARO

Tool use (web fetch, filesystem, time) requires Node.js.

Docker (NVIDIA GPU)

The following commands map the local cache directory to the container in order to persist kernel cache across runs. Remove -v ... to disable this behaviour.

# Interactive chat
docker run --pull=always --rm -it --gpus all --ipc=host \
  -v $HOME/.cache/paroquant:/root/.cache/paroquant \
  ghcr.io/z-lab/paroquant:chat --model z-lab/Qwen3.5-2B-PARO

# API server (port 8000)
docker run --pull=always --rm -it --gpus all --ipc=host -p 8000:8000 \
  -v $HOME/.cache/paroquant:/root/.cache/paroquant \
  ghcr.io/z-lab/paroquant:serve --model z-lab/Qwen3.5-2B-PARO

Citation

@inproceedings{liang2026paroquant,
  title     = {{ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference}},
  author    = {Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026}
}