Xuandong's picture
Improve model card metadata and add paper/code links (#1)
0782f14
|
Raw
History Blame Contribute Delete
1.27 kB
---
base_model:
- allenai/OLMo-2-1124-7B-SFT
datasets:
- math
language:
- en
license: apache-2.0
metrics:
- accuracy
pipeline_tag: text-generation
library_name: transformers
---
# OLMo-2-7B-SFT-GRPO-MATH-1EPOCH-SYSP
**Description:**
A GRPO-fine-tuned version of [allenai/OLMo-2-1124-7B-SFT](https://huggingface.co/allenai/OLMo-2-1124-7B-SFT) trained on the MATH dataset with a system prompt.
This model was developed as part of the research presented in the paper [Learning to Reason without External Rewards](https://huggingface.co/papers/2505.19590). It utilizes the **Intuitor** method, an instantiation of Reinforcement Learning from Internal Feedback (RLIF), which enables models to learn from intrinsic signals like self-certainty without requiring external rewards or labeled gold solutions.
## Resources
- **Paper:** [Learning to Reason without External Rewards](https://huggingface.co/papers/2505.19590)
- **Repository:** [sunblaze-ucb/Intuitor](https://github.com/sunblaze-ucb/Intuitor)
---
## Citation
```bibtex
@article{zhao2025learning,
title={Learning to Reason without External Rewards},
author={Zhao, Xuandong and Kang, Zhewei and Feng, Aosong and Levine, Sergey and Song, Dawn},
journal={arXiv preprint arXiv:2505.19590},
year={2025}
}
```