File size: 539 Bytes
0fe33a2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
---
library_name: transformers
base_model:
- Qwen/Qwen3-4B-Thinking-2507
base_model_relation: quantized
pipeline_tag: text-generation
tags:
- Qwen3
- GPTQ
- W4A16
- llm-compressor
---

# Qwen3-4B-Thinking-2507-llmc-gptq-w4a16-g128-n256-s1024

Base model: `Qwen/Qwen3-4B-Thinking-2507`

Quantized with `llm-compressor`.

- method: `gptq`
- weight format: `W4A16`
- group size: `128`
- calibration dataset: `nvidia/Llama-Nemotron-Post-Training-Dataset`
- calibration split: `math`
- calibration samples: `256`
- max sequence length: `1024`