Instructions to use mwael399/arabic-english-translator with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mwael399/arabic-english-translator with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="mwael399/arabic-english-translator")# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("mwael399/arabic-english-translator") model = AutoModelForSeq2SeqLM.from_pretrained("mwael399/arabic-english-translator", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Arabic-English Translator
ูู
ูุฐุฌ ุชุฑุฌู
ุฉ ู
ู ุงูุนุฑุจูุฉ ุฅูู ุงูุฅูุฌููุฒูุฉุ ุชู
ุจูุงุคู ุนู ุทุฑูู fine-tuning ูู
ูุฏูู Helsinki-NLP/opus-mt-ar-en ุนูู ุฌุฒุก ู
ู dataset opus-100 (ar-en).
A fine-tuned Seq2Seq model for translating Arabic text into English, built on top of
Helsinki-NLP/opus-mt-ar-en and fine-tuned on a subset of the opus-100 (ar-en) dataset.
Model Details
Model Description
ูุฐุง ุงูู ูุฏูู ุฌุฒุก ู ู ู ุดุฑูุน ุชุนููู ู ุดุฎุตู ูุจูุงุก ูุธุงู NLP ุนุฑุจูุ ุชู ุชุฏุฑูุจู ุจุงุณุชุฎุฏุงู ู ูุชุจุฉ ๐ค Transformers ูุชุทุจูู ุนู ูู ุนูู fine-tuning ูู ูุฏููุงุช ุงูุชุฑุฌู ุฉ (Seq2Seq).
- Developed by: Mohamed Wael
- Model type: Seq2Seq Translation (MarianMT architecture)
- Language(s) (NLP): Arabic (ar) โ English (en)
- License: Apache 2.0 (inherited from base model)
- Finetuned from model: Helsinki-NLP/opus-mt-ar-en
Model Sources
- Repository: ู ุชุฑุฌู ุนุฑุจู - ุฅูุฌููุฒู ุนูู GitHub
- Demo: Streamlit App
Uses
Direct Use
ุงูู ูุฏูู ู ูุงุณุจ ูุชุฑุฌู ุฉ ุฌู ู ุนุฑุจูุฉ ุนุงู ุฉ ุฃู ุดุจู ูุตุญู ุฅูู ุงูุฅูุฌููุฒูุฉ ุจุดูู ู ุจุงุดุฑุ ู ู ุบูุฑ ุงูุญุงุฌุฉ ูุฃู ู ุนุงูุฌุฉ ุฅุถุงููุฉ.
Out-of-Scope Use
- ุงูููุฌุฉ ุงูุนุงู ูุฉ: ุงูู ูุฏูู ููุงุฌู ุตุนูุจุฉ ู ุน ุงูููุฌุงุช ุงูุนุงู ูุฉ (ู ุซู ุงูู ุตุฑูุฉ)ุ ุญูุซ ูู ูุชู ุชุฏุฑูุจู ุจุดูู ู ูุซู ุนูู ูุตูุต ุนุงู ูุฉ.
- ุงูู ุตุทูุญุงุช ุงูุชูููุฉ ุงูู ุชุฎุตุตุฉ: ุงูุฃุฏุงุก ูุถุนู ู ุน ุงูู ุตุทูุญุงุช ุงูุชูููุฉ ุงูุฏูููุฉ (ู ุซู ู ุตุทูุญุงุช ุตูุงูุฉ ุงูุณูุงุฑุงุช) ุงูุชู ูุง ุชุธูุฑ ุจูุซุฑุฉ ูู ุจูุงูุงุช ุงูุชุฏุฑูุจ ุงูุนุงู ุฉ (opus-100).
- ุบูุฑ ู ุฎุตุต ููุงุณุชุฎุฏุงู ูู ุชุทุจููุงุช ุทุจูุฉ ุฃู ูุงููููุฉ ุฃู ุฃู ู ุฌุงู ูุชุทูุจ ุฏูุฉ ุชุฑุฌู ุฉ ุญุฑุฌุฉ.
Bias, Risks, and Limitations
- ุงูู ูุฏูู ุชู ุชุฏุฑูุจู ุนูู ุนุฏุฏ ู ุญุฏูุฏ ู ู ุงูุนููุงุช (subset ู ู opus-100)ุ ูููุณ ุนูู ุงูุฏุงุชุงุณุช ุจุงููุงู ูุ ูุฐูู ุฌูุฏุชู ุฃูู ู ู ูู ุงุฐุฌ ุงูุชุฑุฌู ุฉ ุงููุจูุฑุฉ ุงูู ุฏุฑุจุฉ ุนูู ุจูุงูุงุช ุถุฎู ุฉ.
- ูู ุง ูู ุงูุญุงู ู ุน ู ุนุธู ูู ุงุฐุฌ ุงูุชุฑุฌู ุฉ ุงูุขููุฉุ ูุฏ ุชุธูุฑ ุฃุฎุทุงุก ู ุน ุงูุฌู ู ุงูุทูููุฉ ุฃู ุงูู ุนูุฏุฉ ุฃู ุงูุชู ุชุญุชูู ุนูู ุณูุงู ุซูุงูู/ูุบูู ุฎุงุต.
Recommendations
ูููุตุญ ุจุงุณุชุฎุฏุงู ูุฐุง ุงูู ูุฏูู ูุฃุบุฑุงุถ ุชุนููู ูุฉ ุฃู ุชุฌุฑูุจูุฉ. ููุงุณุชุฎุฏุงู ุงูุฅูุชุงุฌู ุงููุนูู ูู ู ุฌุงู ู ุชุฎุตุต (ู ุซู ุตูุงูุฉ ุงูุณูุงุฑุงุช)ุ ูููุตุญ ุจุนู ู fine-tuning ุฅุถุงูู ุนูู ุจูุงูุงุช domain-specific.
How to Get Started with the Model
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("mwael399/arabic-english-translator")
model = AutoModelForSeq2SeqLM.from_pretrained("mwael399/arabic-english-translator")
text = "ู
ุฑุญุจุง ููู ุญุงูู"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Training Details
Training Data
ุฌุฒุก (subset) ู ู dataset opus-100 (ar-en):
- 5000 ุนููุฉ ุชุฏุฑูุจ (train)
- 500 ุนููุฉ ุชูููู (validation)
Training Procedure
Preprocessing
- ุงูุญุฏ ุงูุฃูุตู ูุทูู ุงูู ุฏุฎูุงุช: 128 ุชููู
- ุงูุญุฏ ุงูุฃูุตู ูุทูู ุงูู ุฎุฑุฌุงุช: 128 ุชููู
- ุชู
ุงุณุชุฎุฏุงู
DataCollatorForSeq2Seqููู dynamic padding
Training Hyperparameters
- Learning rate: 2e-5
- Batch size (train): 8
- Batch size (eval): 8
- Epochs: 1
- Weight decay: 0.01
- Training regime: fp32
Evaluation
ูู ูุชู ุฅุฌุฑุงุก ุชูููู ูู ู ุฑุณู ู (ู ุซู BLEU score) ุจุนุฏ ุนูู ู ุฌู ูุนุฉ ุงุฎุชุจุงุฑ ู ููุตูุฉ. ุงูุชูููู ุงูุญุงูู ุชู ุจุดูู ูุฏูู (qualitative) ุนู ุทุฑูู ุงุฎุชุจุงุฑ ุฌู ู ู ุชููุนุฉ.
Summary
ุงูู ูุฏูู ูุนู ู ุจุดูู ุฌูุฏ ู ุน ุงูุฌู ู ุงููุตุญู ุฃู ุดุจู ุงููุตุญู ุงูุจุณูุทุฉุ ูููู ูุธูุฑ ุถุนููุง ูุงุถุญูุง ู ุน ุงูููุฌุฉ ุงูุนุงู ูุฉ ูุงูู ุตุทูุญุงุช ุงูุชูููุฉ ุงูู ุชุฎุตุตุฉุ ูู ุง ูู ู ุชููุน ู ู ุทุจูุนุฉ ุจูุงูุงุช ุงูุชุฏุฑูุจ.
Technical Specifications
Model Architecture and Objective
MarianMT (Seq2Seq Transformer) architecture, ู
ุจูู ุฃุณุงุณูุง ุนูู Helsinki-NLP/opus-mt-ar-en.
Compute Infrastructure
Software
- ๐ค Transformers
- PyTorch
- ๐ค Datasets
Model Card Authors
Mohamed Wael
Model Card Contact
ููุชูุงุตู: ุนุจุฑ Hugging Face profile mwael399
- Downloads last month
- -
Model tree for mwael399/arabic-english-translator
Base model
Helsinki-NLP/opus-mt-ar-en