qwen3-32B-baseline-iter139
GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page /
finish tools over a fixed 100K-doc corpus). Cell qwen3-32B/baseline of the
length-penalty experiment matrix, checkpoint at training iteration 139.
- Code: https://github.com/ys-2020/miles (branch
browsecomp-rl, seedocs/experiments/browsecomp-length-penalty-results.mdfor the full study) - WandB: project
browsecomp-b300 - Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.
Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)
| iter | accuracy | mean response len (tokens) | truncated ratio |
|---|---|---|---|
| 19 | 0.293 | 2482 | 0.24 |
| 39 | 0.313 | 2586 | 0.19 |
| 59 | 0.347 | 2597 | 0.18 |
| 79 | 0.400 | 2527 | 0.13 |
| 99 | 0.407 | 2803 | 0.15 |
| 119 | 0.387 | 2820 | 0.13 |
| 139 | 0.427 | 3102 | 0.09 |
Resuming training in miles
Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to
iter_0000139, write 139 into latest_checkpointed_iteration.txt, then
launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).
training_metadata.json in this repo records provenance.
- Downloads last month
- 22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for Shangy/browsecomp-qwen3-32B-baseline-iter139
Base model
Qwen/Qwen3-32B