Qwen3-30B-A3B midtrain (swe50-star50)
Continued pretraining ("midtrain") of Qwen/Qwen3-30B-A3B-Base on the same 50/50 mixture as
GLM-4.5-Air-midtrain-swe50-star50-90B:
SWE-agent-trajectory tokens + StarCoder-v2-anchored code/math/STEM/web text. 52k sequence length,
trained to iter 19,267 (final; wandb kr8fvpmh). Converted from Megatron torch-dist to HF safetensors.
Note: on this logit-distilled Qwen base, midtrain RAISED aggregate val loss while still helping downstream agentic benchmarks after SFT -- see the ablation curves in openrecipe-configs-and-ablations. Ratio-ablation siblings (star100 / star80-swe20 / star20-swe80 / star10-swe90 / swe100) exist as torch-dist checkpoints; configs and wandb curves are in the same dataset repo.
- Midtrain author: Wai Tong Chung. Released as part of a Duke University research project.
- Downloads last month
- 10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for TianchenGuan/Qwen3-30B-A3B-midtrain-swe50-star50
Base model
Qwen/Qwen3-30B-A3B-Base