Qwen3-4B-Instruct-2507 — tau2-bench telecom, final_pairs_2 (120)
Qwen3-4B-Instruct-2507 consolidated with OEL on the tau2-bench telecom domain, using
120 externally authored task–memory pairs from the revised final_pairs_2 set,
stratified down from 125 so the batches divide evenly.
Sibling runs: telecom-fp84-1ep, telecom-compsub-1ep, telecom-ours-1ep, telecom-oel-1ep.
Training
| base model | Qwen3-4B-Instruct-2507 |
| training pairs | 120 (composition 56 / recall 38 / substitution 26) |
| batch / steps | 12 / 10 (1 epoch, every batch full) |
| learning rate | 3e-6, constant |
| gradient clipping | 1.0 (verl default) |
| loss | full KL over all response tokens, kl_topk 256 |
| user simulator | gpt-4.1-mini (temperature 0) |
The student replays each task without memory; the teacher is the same weights with that task's memory in the system prompt. Only the prompt differs, and no reward is used. Five pairs were dropped from the 125 by stratified sampling (seed 42), holding the level ratios to within 1pp, so that no batch is short — earlier runs showed a tiny final batch wrecks the saved checkpoint, and only the last step is saved.
Training trace
| step | KL loss | entropy | grad norm |
|---|---|---|---|
| 1 | 0.024 | 0.341 | 5.654 |
| 2 | 0.015 | 0.318 | 1.687 |
| 3 | 0.014 | 0.307 | 12.898 |
| 4 | 0.019 | 0.474 | 11.900 |
| 5 | 0.054 | 0.385 | 7.949 |
| 6 | 0.023 | 0.362 | 3.851 |
| 7 | 0.020 | 0.394 | 11.133 |
| 8 | 0.021 | 0.408 | 6.128 |
| 9 | 0.049 | 0.345 | 7.914 |
| 10 | 0.034 | 0.347 | 4.820 |
Each step reads a different batch, so the loss column tracks batch difficulty rather than convergence. The gradient norm sits above the 1.0 clip on every step — telecom episodes are long, so the effective step size is set by the clip rather than by the learning rate.
Evaluation
Not evaluated. Pushed straight after training. On tau2-bench telecom test (40 tasks,
4 trials, no memory at eval) the base model scores avg 0.056 / pass@4 0.175; other runs
in this series land between 0.050 and 0.081, all within about 1.5 SE of each other.
Known caveat. A substitution pair in this family reproduced a test task's exact
environment state by dropping a fault from a training task; task ids still carry the
training signature, so only an initial_state comparison reveals it.
Usage
vllm serve MANGSEOK123/qwen3-4b-tau2-telecom-fp120-1ep \
--enable-auto-tool-choice --tool-call-parser hermes --max-model-len 40960
- Downloads last month
- 28
Model tree for MANGSEOK123/qwen3-4b-tau2-telecom-fp120-1ep
Base model
Qwen/Qwen3-4B-Instruct-2507