Trained ChatR1 model on TopiOCQA with PPO. See ChatR1 paper for reference (https://arxiv.org/abs/2510.13312). ```bibtex @article{lupart2025chatr1, title={Chatr1: Reinforcement learning for conversational reasoning and retrieval augmented question answering}, author={Lupart, Simon and Aliannejadi, Mohammad and Kanoulas, Evangelos}, journal={arXiv preprint arXiv:2510.13312}, year={2025} } ```