Rubin-Wei's picture
Update README.md
f3b2484 verified
|
Raw
History Blame Contribute Delete
1.5 kB
metadata
license: apache-2.0
language:
  - en
base_model:
  - openai-community/gpt2-xl
datasets:
  - wikitext-103

GPT2-large-Finetuned-WikiText103

Model Description

A fine-tuned version of GPT2-large on the WikiText-103 dataset.

Performance on WikiText-103

Model Perplexity Improvement
GPT2-large (baseline) 15.80 -
GPT2-large-Finetuned 10.42 -5.38

Training Details

  • Training Data: WikiText-103 (103M tokens)
  • Optimizer: AdamW
  • Learning Rate: 2e-5 with cosine schedule

Citation

This model was released as part of the paper "MLP Memory: A Retriever-Pretrained Memory for Large Language Models".

For more information, see: https://github.com/Binn0/MLPMemory.

If you use this model, please cite:

@inproceedings{Wei2025MLPMA,
  title={MLP Memory: A Retriever-Pretrained Memory for Large Language Models},
  author={Rubin Wei and Jiaqi Cao and Jiarui Wang and Jushi Kai and Qipeng Guo and Bowen Zhou and Zhouhan Lin},
  year={2025},
  url={https://api.semanticscholar.org/CorpusID:281658735}
}