Qwen3-4B-Instruct-2507 — tau2-bench telecom, final_pairs_2 (120)

Qwen3-4B-Instruct-2507 consolidated with OEL on the tau2-bench telecom domain, using 120 externally authored task–memory pairs from the revised final_pairs_2 set, stratified down from 125 so the batches divide evenly.

Sibling runs: telecom-fp84-1ep, telecom-compsub-1ep, telecom-ours-1ep, telecom-oel-1ep.

Training

base model Qwen3-4B-Instruct-2507
training pairs 120 (composition 56 / recall 38 / substitution 26)
batch / steps 12 / 10 (1 epoch, every batch full)
learning rate 3e-6, constant
gradient clipping 1.0 (verl default)
loss full KL over all response tokens, kl_topk 256
user simulator gpt-4.1-mini (temperature 0)

The student replays each task without memory; the teacher is the same weights with that task's memory in the system prompt. Only the prompt differs, and no reward is used. Five pairs were dropped from the 125 by stratified sampling (seed 42), holding the level ratios to within 1pp, so that no batch is short — earlier runs showed a tiny final batch wrecks the saved checkpoint, and only the last step is saved.

Training trace

step KL loss entropy grad norm
1 0.024 0.341 5.654
2 0.015 0.318 1.687
3 0.014 0.307 12.898
4 0.019 0.474 11.900
5 0.054 0.385 7.949
6 0.023 0.362 3.851
7 0.020 0.394 11.133
8 0.021 0.408 6.128
9 0.049 0.345 7.914
10 0.034 0.347 4.820

Each step reads a different batch, so the loss column tracks batch difficulty rather than convergence. The gradient norm sits above the 1.0 clip on every step — telecom episodes are long, so the effective step size is set by the clip rather than by the learning rate.

Evaluation

Not evaluated. Pushed straight after training. On tau2-bench telecom test (40 tasks, 4 trials, no memory at eval) the base model scores avg 0.056 / pass@4 0.175; other runs in this series land between 0.050 and 0.081, all within about 1.5 SE of each other.

Known caveat. A substitution pair in this family reproduced a test task's exact environment state by dropping a fault from a training task; task ids still carry the training signature, so only an initial_state comparison reveals it.

Usage

vllm serve MANGSEOK123/qwen3-4b-tau2-telecom-fp120-1ep \
  --enable-auto-tool-choice --tool-call-parser hermes --max-model-len 40960
Downloads last month
28
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MANGSEOK123/qwen3-4b-tau2-telecom-fp120-1ep

Finetuned
(2375)
this model