Qwen3-4B-Instruct-2507 — tau2-bench telecom, final_pairs (84)

Qwen3-4B-Instruct-2507 consolidated with OEL on the tau2-bench telecom domain, using 84 externally authored task–memory pairs — a revised final_pairs set, stratified down from 85 so the batches divide evenly.

Sibling runs on the same domain and protocol: telecom-compsub-1ep, telecom-ours-1ep, telecom-oel-1ep.

Training

Hyperparameters copied from the retail memory_aug_final run, the one configuration in this series that scored above its base model, so the dataset is the only thing that differs.

base model Qwen3-4B-Instruct-2507
training pairs 84 (composition 46 / recall 19 / substitution 19)
batch / steps 12 / 7 (1 epoch, every batch full)
learning rate 3e-6, constant
gradient clipping 1.0 (verl default)
loss full KL over all response tokens, kl_topk 256
user simulator gpt-4.1-mini (temperature 0)

The student replays each task without memory; the teacher is the same weights with that task's memory in the system prompt. Only the prompt differs, and no reward is used.

The source set had 85 pairs, which leaves a final batch of 1 at batch size 12. Earlier runs in this series showed that a tiny last batch wrecks the saved checkpoint — entropy collapsed to 0.042 on a 4-sample final step, and only the last step is saved. One pair was dropped by stratified sampling (seed 42), holding the level ratios to within 1pp.

Training trace

step KL loss entropy grad norm
1 0.022 0.309 9.54
2 0.011 0.247 4.83
3 0.016 0.301 6.52
4 0.025 0.343 6.28
5 0.032 0.272 7.67
6 0.016 0.337 5.56
7 0.029 0.336 7.75

Each step reads a different batch, so the loss column tracks batch difficulty rather than convergence. Entropy stays in a healthy band through the final step. Note the gradient norm sits above the 1.0 clip on every step — telecom episodes are long, so the effective step size is set by the clip rather than by the learning rate.

Evaluation

Not evaluated. Pushed straight after training. For context, on tau2-bench telecom test (40 tasks, 4 trials, no memory at eval) the base model scores avg 0.056 / pass@4 0.175; other runs in this series land between 0.050 and 0.081, all within about 1.5 SE of each other.

Known caveat. A substitution pair in this family reproduced a test task's exact environment state (break_apn_settings + lock_sim_card_pin) by dropping a fault from a training task. Task ids still carry the training signature, so only an initial_state comparison reveals it.

Usage

vllm serve MANGSEOK123/qwen3-4b-tau2-telecom-fp84-1ep \
  --enable-auto-tool-choice --tool-call-parser hermes --max-model-len 40960
Downloads last month
24
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MANGSEOK123/qwen3-4b-tau2-telecom-fp84-1ep

Finetuned
(2375)
this model