--- license: gemma base_model: google/translategemma-4b-it language: - ar - de pipeline_tag: translation tags: - translation - arabic - german - islam - khutbah - quran - gemma3 - qlora --- # TranslateGemma-4B-Khutbah Arabic → German translation model specialized for **Islamic sermons (khutbahs)** — the compact sibling of [Yacinedh/translategemma-12b-khutbah](https://huggingface.co/Yacinedh/translategemma-12b-khutbah). Fine-tuned from [google/translategemma-4b-it](https://huggingface.co/google/translategemma-4b-it). **Headline: after domain fine-tuning, this 4B matches the stock 12B** on the khutbah benchmark at 40% of the size — making it a strong choice for CPU/edge deployment (e.g. an offline fallback on a mosque laptop). ## Results 74-case khutbah benchmark (embedding cosine vs held-out references, identical pipeline for all rows): | Model | Overall | Free sermon rhetoric | |---|---|---| | google/translategemma-4b-it (stock) | 0.9368 | 0.8588 | | **this model** | **0.9527** | **0.8789** | | google/translategemma-12b-it (stock, reference) | 0.9535 | 0.8836 | ## Training Same recipe and data as the 12B: 24,240 Arabic–German pairs (four German Quran editions, liturgical formulas, hadith, terminology; benchmark sentences excluded), QLoRA r=16, lr 1e-4, 1 epoch. Trained on a single T4 (fp16). Code: [MinbarAI/training](https://github.com/Yacine-DH/MinbarAI/tree/main/training). ## Files - Merged fp16 safetensors - `model-Q4_K_M.gguf` (~2.5 GB) — runs on CPU via llama.cpp / Ollama ## Usage Identical to the 12B card — TranslateGemma structured chat template (`source_lang_code: "ar"`, `target_lang_code: "de-DE"`). See [Yacinedh/translategemma-12b-khutbah](https://huggingface.co/Yacinedh/translategemma-12b-khutbah) for snippets. ## License Gemma Terms of Use. Quran translation data from [Tanzil](https://tanzil.net/trans/) via [fawazahmed0/quran-api](https://github.com/fawazahmed0/quran-api).