allenai/tulu-3-sft-mixture
Viewer • Updated • 939k • 35.2k • 255
Short-run supervised finetune of allenai/Olmo-3-1025-7B on allenai/tulu-3-sft-mixture for 1000 optimizer steps (~7% of one epoch). Used as a "warm teacher" in ambient knowledge distillation experiments — partially instruction-tuned, sitting between pretrained base and a fully SFT-d model.
allenai/Olmo-3-1025-7B (pretrained)allenai/tulu-3-sft-mixture (~939k examples)--add_bos, seed 123Designed as a warm teacher for ambient knowledge distillation. The Llama-3.1-8B sibling (giannisdaras/tulu-vista-repro is the fully-SFT-d 14613-step Llama; the Llama warm teacher used a separate 1k Llama checkpoint) shows that 1k-step warm teachers can give modest gains over base-model cold teachers.
For instruction following, use a fully SFT-d Olmo3-Tulu instead.
Base model
allenai/Olmo-3-1025-7B