Qwen3-4B-Instruct-2507 — tau2-bench telecom, final_pairs (84)
Qwen3-4B-Instruct-2507 consolidated with OEL on the tau2-bench telecom domain, using
84 externally authored task–memory pairs — a revised final_pairs set, stratified down
from 85 so the batches divide evenly.
Sibling runs on the same domain and protocol: telecom-compsub-1ep, telecom-ours-1ep, telecom-oel-1ep.
Training
Hyperparameters copied from the retail memory_aug_final run, the one configuration in
this series that scored above its base model, so the dataset is the only thing that differs.
| base model | Qwen3-4B-Instruct-2507 |
| training pairs | 84 (composition 46 / recall 19 / substitution 19) |
| batch / steps | 12 / 7 (1 epoch, every batch full) |
| learning rate | 3e-6, constant |
| gradient clipping | 1.0 (verl default) |
| loss | full KL over all response tokens, kl_topk 256 |
| user simulator | gpt-4.1-mini (temperature 0) |
The student replays each task without memory; the teacher is the same weights with that task's memory in the system prompt. Only the prompt differs, and no reward is used.
The source set had 85 pairs, which leaves a final batch of 1 at batch size 12. Earlier runs in this series showed that a tiny last batch wrecks the saved checkpoint — entropy collapsed to 0.042 on a 4-sample final step, and only the last step is saved. One pair was dropped by stratified sampling (seed 42), holding the level ratios to within 1pp.
Training trace
| step | KL loss | entropy | grad norm |
|---|---|---|---|
| 1 | 0.022 | 0.309 | 9.54 |
| 2 | 0.011 | 0.247 | 4.83 |
| 3 | 0.016 | 0.301 | 6.52 |
| 4 | 0.025 | 0.343 | 6.28 |
| 5 | 0.032 | 0.272 | 7.67 |
| 6 | 0.016 | 0.337 | 5.56 |
| 7 | 0.029 | 0.336 | 7.75 |
Each step reads a different batch, so the loss column tracks batch difficulty rather than convergence. Entropy stays in a healthy band through the final step. Note the gradient norm sits above the 1.0 clip on every step — telecom episodes are long, so the effective step size is set by the clip rather than by the learning rate.
Evaluation
Not evaluated. Pushed straight after training. For context, on tau2-bench telecom
test (40 tasks, 4 trials, no memory at eval) the base model scores avg 0.056 / pass@4 0.175; other runs in this series land between 0.050 and 0.081, all within about
1.5 SE of each other.
Known caveat. A substitution pair in this family reproduced a test task's exact
environment state (break_apn_settings + lock_sim_card_pin) by dropping a fault from a
training task. Task ids still carry the training signature, so only an initial_state
comparison reveals it.
Usage
vllm serve MANGSEOK123/qwen3-4b-tau2-telecom-fp84-1ep \
--enable-auto-tool-choice --tool-call-parser hermes --max-model-len 40960
- Downloads last month
- 24
Model tree for MANGSEOK123/qwen3-4b-tau2-telecom-fp84-1ep
Base model
Qwen/Qwen3-4B-Instruct-2507