YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

๐Ÿš€ English โ†’ Hindi Transformer

Python PyTorch Task BLEU Optimization


๐Ÿ“Œ Overview

This model implements a Transformer-based Neural Machine Translation (NMT) system for English โ†’ Hindi translation using PyTorch.

Optimized using Ray Tune + Optuna + ASHA, achieving high BLEU score with reduced training time.


๐Ÿ—๏ธ Model Architecture

  • Encoderโ€“Decoder Transformer
  • 6 Encoder + 6 Decoder layers
  • Multi-head attention
  • Feed-forward network
  • Positional encoding
  • Residual connections + LayerNorm

Configuration

  • d_model: 512
  • num_heads: 4
  • num_layers: 6
  • d_ff: 4096
  • dropout: 0.054

โš™๏ธ Training Details

  • Dataset: ~13,186 English-Hindi sentence pairs
  • Optimizer: AdamW
  • Loss: CrossEntropy (ignore padding)
  • Scheduler: CosineAnnealingLR
  • Device: GPU

๐Ÿ“Š Results

Metric Baseline Tuned
Epochs 100 30
Time 129.42 min 79.31 min
Loss 0.0972 0.0959
BLEU 68.02 90.38

โš™๏ธ Best Hyperparameters

  • LR: 0.0001009
  • Batch size: 64
  • Heads: 4
  • d_ff: 4096
  • Dropout: 0.054
  • Weight decay: 0.000261

๐Ÿงช Sample Outputs

  • EN: I love you HI: เคฎเฅˆเค‚ เคคเฅเคฎเคธเฅ‡ เคชเฅเคฏเคพเคฐ เค•เคฐเคคเคพ เคนเฅ‚เค

  • EN: What is your name? HI: เค†เคชเค•เคพ เคจเคพเคฎ เค•เฅเคฏเคพ เคนเฅˆ?

  • EN: How are you? HI: เค†เคช เค•เฅˆเคธเฅ‡ เคนเฅˆเค‚?


๐Ÿ“‚ Files

  • M25CSA011_ass_4_best_model.pth
  • en_vocab.pkl
  • hi_vocab.pkl
  • best_config.json

๐Ÿ‘ค Author

Mahek Shankesh Gadiya M.Tech AI โ€“ IIT Jodhpur


๐Ÿ“š Assignment

Transformer Optimization using Ray Tune + Optuna

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support