qwen3-8B-length-penalty-iter199
GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page /
finish tools over a fixed 100K-doc corpus). Cell qwen3-8B/length_penalty of the
length-penalty experiment matrix, checkpoint at training iteration 199.
- Code: https://github.com/ys-2020/miles (branch
browsecomp-rl, seedocs/experiments/browsecomp-length-penalty-results.mdfor the full study) - WandB: project
browsecomp-b300 - Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.
Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)
| iter | accuracy | mean response len (tokens) | truncated ratio |
|---|---|---|---|
| 19 | 0.260 | 1973 | 0.25 |
| 39 | 0.240 | 2237 | 0.29 |
| 59 | 0.293 | 2080 | 0.30 |
| 79 | 0.280 | 2139 | 0.25 |
| 99 | 0.240 | 2156 | 0.31 |
| 119 | 0.300 | 2463 | 0.35 |
| 139 | 0.320 | 2521 | 0.36 |
| 159 | 0.353 | 2466 | 0.23 |
| 179 | 0.340 | 2563 | 0.19 |
| 199 | 0.373 | 2405 | 0.15 |
Resuming training in miles
Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to
iter_0000199, write 199 into latest_checkpointed_iteration.txt, then
launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).
training_metadata.json in this repo records provenance.
- Downloads last month
- 26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support