YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
๐ English โ Hindi Transformer
๐ Overview
This model implements a Transformer-based Neural Machine Translation (NMT) system for English โ Hindi translation using PyTorch.
Optimized using Ray Tune + Optuna + ASHA, achieving high BLEU score with reduced training time.
๐๏ธ Model Architecture
- EncoderโDecoder Transformer
- 6 Encoder + 6 Decoder layers
- Multi-head attention
- Feed-forward network
- Positional encoding
- Residual connections + LayerNorm
Configuration
- d_model: 512
- num_heads: 4
- num_layers: 6
- d_ff: 4096
- dropout: 0.054
โ๏ธ Training Details
- Dataset: ~13,186 English-Hindi sentence pairs
- Optimizer: AdamW
- Loss: CrossEntropy (ignore padding)
- Scheduler: CosineAnnealingLR
- Device: GPU
๐ Results
| Metric | Baseline | Tuned |
|---|---|---|
| Epochs | 100 | 30 |
| Time | 129.42 min | 79.31 min |
| Loss | 0.0972 | 0.0959 |
| BLEU | 68.02 | 90.38 |
โ๏ธ Best Hyperparameters
- LR: 0.0001009
- Batch size: 64
- Heads: 4
- d_ff: 4096
- Dropout: 0.054
- Weight decay: 0.000261
๐งช Sample Outputs
EN: I love you HI: เคฎเฅเค เคคเฅเคฎเคธเฅ เคชเฅเคฏเคพเคฐ เคเคฐเคคเคพ เคนเฅเค
EN: What is your name? HI: เคเคชเคเคพ เคจเคพเคฎ เคเฅเคฏเคพ เคนเฅ?
EN: How are you? HI: เคเคช เคเฅเคธเฅ เคนเฅเค?
๐ Files
- M25CSA011_ass_4_best_model.pth
- en_vocab.pkl
- hi_vocab.pkl
- best_config.json
๐ค Author
Mahek Shankesh Gadiya M.Tech AI โ IIT Jodhpur
๐ Assignment
Transformer Optimization using Ray Tune + Optuna
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support