trillionlabs/SimScholar-RL
Viewer • Updated • 13k • 39
Datasets and model checkpoints for training and studying scientific-literature search agents in the S3 local environment.
Note 13,000 synthetic single-hop literature-search QA items for agentic reinforcement learning.
Note 14,633 labeled single- and two-hop ReAct trajectories with complete tool-use conversations.
Note Qwen3-4B research checkpoint trained with agentic RL in the S3 literature-search environment.