Qwen3-4B-Thinking-2507-llmc-awq-calib-chat-w4a16-g128-n256-s1024

Base model: Qwen/Qwen3-4B-Thinking-2507

Quantized with llm-compressor.

  • method: awq
  • weight format: W4A16
  • group size: 128
  • calibration dataset: nvidia/Llama-Nemotron-Post-Training-Dataset
  • calibration split: chat
  • calibration samples: 256
  • max sequence length: 1024
Downloads last month
8
Safetensors
Model size
4B params
Tensor type
I64
I32
BF16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for meteorain/Qwen3-4B-Thinking-2507-llmc-awq-calib-chat-w4a16-g128-n256-s1024

Quantized
(136)
this model