File size: 1,940 Bytes
6a39653
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
---
language:
- en
license: apache-2.0
library_name: gguf
tags:
- reranker
- gguf
- llama.cpp
base_model: Qwen/Qwen3-Reranker-8B
---

# Qwen3-Reranker-8B-Q4_K_M-GGUF

This model was converted to GGUF format from [Qwen/Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B) using llama.cpp via the ggml.ai's GGUF-my-repo space.

Refer to the [original model card](https://huggingface.co/Qwen/Qwen3-Reranker-8B) for more details on the model.

## Model Information

- **Base Model**: [Qwen/Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B)
- **Quantization**: Q4_K_M
- **Format**: GGUF (GPT-Generated Unified Format)
- **Converted with**: llama.cpp

## Quantization Details

This is a **Q4_K_M** quantization of the original model:

- **F16**: Full 16-bit floating point - highest quality, largest size
- **Q8_0**: 8-bit quantization - high quality, good balance
- **Q4_K_M**: 4-bit quantization with medium quality - smaller size, faster inference

## Usage

This model can be used with llama.cpp and other GGUF-compatible inference engines.

```bash
# Example using llama.cpp
./llama-rerank -m Qwen3-Reranker-8B-Q4_K_M.gguf
```

## Model Files

| Quantization | Use Case |
|-------------|----------|
| F16 | Maximum quality, largest size |
| Q8_0 | High quality, good balance of size/performance |
| Q4_K_M | Good quality, smallest size, fastest inference |

## Citation

If you use this model, please cite the original model:

```bibtex
# See original model card for citation information
```

## License

This model inherits the license from the original model. Please refer to the [original model card](https://huggingface.co/Qwen/Qwen3-Reranker-8B) for license details.

## Acknowledgements

- Original model by the authors of [Qwen/Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B)
- GGUF conversion via llama.cpp by ggml.ai
- Converted and uploaded by [sinjab](https://huggingface.co/sinjab)