qwen3-32B-baseline-iter139

GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page / finish tools over a fixed 100K-doc corpus). Cell qwen3-32B/baseline of the length-penalty experiment matrix, checkpoint at training iteration 139.

  • Code: https://github.com/ys-2020/miles (branch browsecomp-rl, see docs/experiments/browsecomp-length-penalty-results.md for the full study)
  • WandB: project browsecomp-b300
  • Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.

Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)

iter accuracy mean response len (tokens) truncated ratio
19 0.293 2482 0.24
39 0.313 2586 0.19
59 0.347 2597 0.18
79 0.400 2527 0.13
99 0.407 2803 0.15
119 0.387 2820 0.13
139 0.427 3102 0.09

Resuming training in miles

Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to iter_0000139, write 139 into latest_checkpointed_iteration.txt, then launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).

training_metadata.json in this repo records provenance.

Downloads last month
22
Safetensors
Model size
33B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shangy/browsecomp-qwen3-32B-baseline-iter139

Base model

Qwen/Qwen3-32B
Finetuned
(524)
this model