Qwen3-30B-A3B midtrain (swe50-star50)

Continued pretraining ("midtrain") of Qwen/Qwen3-30B-A3B-Base on the same 50/50 mixture as GLM-4.5-Air-midtrain-swe50-star50-90B: SWE-agent-trajectory tokens + StarCoder-v2-anchored code/math/STEM/web text. 52k sequence length, trained to iter 19,267 (final; wandb kr8fvpmh). Converted from Megatron torch-dist to HF safetensors.

Note: on this logit-distilled Qwen base, midtrain RAISED aggregate val loss while still helping downstream agentic benchmarks after SFT -- see the ablation curves in openrecipe-configs-and-ablations. Ratio-ablation siblings (star100 / star80-swe20 / star20-swe80 / star10-swe90 / swe100) exist as torch-dist checkpoints; configs and wandb curves are in the same dataset repo.

  • Midtrain author: Wai Tong Chung. Released as part of a Duke University research project.
Downloads last month
10
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TianchenGuan/Qwen3-30B-A3B-midtrain-swe50-star50

Finetuned
(63)
this model