---
library_name: transformers
license: apache-2.0
base_model: Qwen/Qwen3-8B
language:
- af # Afrikaans
- am # Amharic
- ar # Arabic
- en # English
- fr # French
- ha # Hausa
- ig # Igbo
- mg # Malagasy (Plateau)
- ny # Nyanja
- om # Oromo
- pt # Portuguese
- rw # Kinyarwanda
- sn # Shona
- so # Somali
- st # Southern Sotho
- sw # Swahili
- ti # Tigrinya
- tn # Tswana
- xh # Xhosa
- yo # Yoruba
- zu # Zulu
pipeline_tag: text-generation
tags:
- african-languages
- multilingual
- continued-pretraining
- afrique-llm
- qwen
---
# AfriqueQwen-8B
## Model Overview
**AfriqueQwen-8B** is part of the **AfriqueLLM** suite—a collection of open language models adapted to **20 African languages** through continued pre-training (CPT) on **26B tokens**. This model is based on [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) and has been specifically adapted for improved performance on African languages while maintaining strong capabilities in high-resource languages.
Our experiments show that **Qwen 3 models achieve the best performance** among all base models tested, better preserving performance in high-resource languages after CPT and achieving strong results on long-context tasks such as document-level translation.
### Key Features
- **Type**: Causal Language Model (Base/Pre-trained)
- **Base Model**: Qwen 3 8B
- **Parameters**: 8B
- **Context Length**: 32,768 tokens (native)
- **Training Tokens**: 26B tokens of carefully curated multilingual data
## Supported Languages
AfriqueQwen-8B has been adapted for the following 20 African languages plus 4 high-resource languages:
| Language | Code | Family | Script |
|----------|------|--------|--------|
| Afrikaans | afr_Latn | Germanic | Latin |
| Swahili | swh_Latn | Bantu | Latin |
| Moroccan Arabic | ary_Arab | Semitic | Arabic |
| Somali | som_Latn | Cushitic | Latin |
| Amharic | amh_Ethi | Semitic | Ethiopic |
| Egyptian Arabic | arz_Arab | Semitic | Arabic |
| Hausa | hau_Latn | Chadic | Latin |
| Kinyarwanda | kin_Latn | Bantu | Latin |
| Zulu | zul_Latn | Bantu | Latin |
| Igbo | ibo_Latn | Volta-Niger | Latin |
| Plateau Malagasy | plt_Latn | Austronesian | Latin |
| Xhosa | xho_Latn | Bantu | Latin |
| Shona | sna_Latn | Bantu | Latin |
| Yoruba | yor_Latn | Volta-Niger | Latin |
| Nyanja | nya_Latn | Bantu | Latin |
| Southern Sotho | sot_Latn | Bantu | Latin |
| Tigrinya | tir_Ethi | Semitic | Ethiopic |
| Tunisian Arabic | aeb_Arab | Semitic | Arabic |
| Oromo | gaz_Latn | Cushitic | Latin |
| Tswana | tsn_Latn | Bantu | Latin |
**High-resource languages (for catastrophic forgetting mitigation):** English, French, Portuguese, Arabic
## Training Data
Our training corpus combines multiple high-quality sources:
- **African Monolingual Data** (~22.8B tokens): FineWeb2, WURA, and MADLAD-400
- **Code** (~1B tokens): CornStack-Python for reasoning capabilities
- **Mathematics** (~1B tokens): FineMath-4+ for mathematical understanding
- **Synthetic Data** (~324M tokens): GPT-4.1 translated domain-specific content across 10 domains
- **Parallel Data** (~456M tokens): NLLB-OPUS filtered with SSA-COMET (threshold 0.7)
We use **UniMax sampling** to create a balanced distribution, capping high-resource languages at approximately 1B tokens and upsampling lower-resource languages for up to five epochs.
## Quickstart
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "McGill-NLP/AfriqueQwen-8B"
# Load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
# Prepare the model input
prompt = "Bawo ni o ṣe n ṣe?" # Yoruba: "How are you doing?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
# Generate text
generated_ids = model.generate(
**inputs,
max_new_tokens=100,
do_sample=True,
temperature=0.7,
top_p=0.9
)
output = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(output)
```
## Deployment
For deployment, you can use `vllm` or `sglang` to create an OpenAI-compatible API endpoint:
**vLLM:**
```shell
vllm serve McGill-NLP/AfriqueQwen-8B
```
**SGLang:**
```shell
python -m sglang.launch_server --model-path McGill-NLP/AfriqueQwen-8B
```
## Training Details
### Hyperparameters
- **Learning Rate**: 5e-5 (with warmup and cosine decay)
- **Context Length**: 8,192 tokens (training)
- **Batch Size**: Effective batch size optimized for throughput
- **Optimizer**: AdamW
- **Precision**: BF16 mixed precision
### Infrastructure
Training was conducted using the LLaMA-Factory framework on up to 64 NVIDIA H100 GPUs with:
- DeepSpeed ZeRO-1/ZeRO-2
- Flash Attention 3
- Sequence packing
- Liger Kernel optimizations
## Evaluation
AfriqueQwen-8B is evaluated on multiple multilingual benchmarks including:
- **AfriMGSM**: Mathematical reasoning
- **AfriMMLU**: Multilingual knowledge
- **AfriXNLI**: Natural language inference
- **Belebele**: Reading comprehension
- **SIB-200**: Topic classification
- **FLORES**: Machine translation
## Model Variants
- [AfriqueQwen-14B](https://huggingface.co/McGill-NLP/AfriqueQwen-14B) - Larger, more capable variant
## Intended Use
This model is designed for:
- Research on African language NLP
- Building applications for African language communities
- Cross-lingual transfer learning experiments
- Multilingual text generation and understanding
## Limitations
- This is a **base/pre-trained model** and may require fine-tuning for specific tasks
- Performance varies across languages based on data availability
- May generate biased or inappropriate content without proper safeguards
- Not suitable for production use without additional safety measures
## Citation
If you find our work helpful, please cite:
```bibtex
@article{afriquellm2025,
title={AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages},
author={AfriqueLLM Team},
year={2025}
}
```
## License
This model is released under the [Apache 2.0 License](https://www.apache.org/licenses/LICENSE-2.0). Please review the license terms before use.
## Acknowledgments
We thank the creators of the base models and datasets that made this work possible, including Alibaba (Qwen), the FineWeb team, WURA, MADLAD-400, and the NLLB project.