qwen3-8B-length-penalty-iter199

GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page / finish tools over a fixed 100K-doc corpus). Cell qwen3-8B/length_penalty of the length-penalty experiment matrix, checkpoint at training iteration 199.

  • Code: https://github.com/ys-2020/miles (branch browsecomp-rl, see docs/experiments/browsecomp-length-penalty-results.md for the full study)
  • WandB: project browsecomp-b300
  • Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.

Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)

iter accuracy mean response len (tokens) truncated ratio
19 0.260 1973 0.25
39 0.240 2237 0.29
59 0.293 2080 0.30
79 0.280 2139 0.25
99 0.240 2156 0.31
119 0.300 2463 0.35
139 0.320 2521 0.36
159 0.353 2466 0.23
179 0.340 2563 0.19
199 0.373 2405 0.15

Resuming training in miles

Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to iter_0000199, write 199 into latest_checkpointed_iteration.txt, then launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).

training_metadata.json in this repo records provenance.

Downloads last month
26
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shangy/browsecomp-qwen3-8B-length-penalty-iter199

Finetuned
Qwen/Qwen3-8B
Finetuned
(1964)
this model