Qwen3-4B-Instruct-2507 โ€” tau2-bench telecom, externally authored memories ("ours")

Qwen3-4B-Instruct-2507 consolidated with OEL on the tau2-bench telecom domain, using memories written by a separate authoring pipeline rather than extracted from the model's own rollouts. The companion model qwen3-4b-tau2-telecom-oel-1ep is the self-extracted counterpart, trained the same way.

How it was made

  1. Deploy โ€” the student plays each of the 100 taskโ€“memory pairs without memory; every agent turn is saved.
  2. Consolidate โ€” the same trajectories are scored again with each task's memory injected into the system prompt (the teacher), and the memory-free student is fit to that distribution with a token-level full KL loss. Teacher and student share weights; the only difference is the prompt. No reward signal is used.
base model Qwen3-4B-Instruct-2507
training pairs 100 (telecom_final_pairs_format.json, 50 source tasks ร— 2 memories)
batch / steps 12 / 9 (1 epoch)
learning rate 3e-6, constant
loss full KL, all response tokens
user simulator gpt-4.1-mini (temperature 0)

Results โ€” tau2-bench telecom test (40 tasks, 4 trials, no memory at eval)

pass@3 pass@4 pass^3 avg reward
base Qwen3-4B-Instruct-2507 0.144 0.175 0.000 0.056
this model (external memories) 0.144 0.175 0.006 0.062
self-extracted, 1 epoch 0.156 0.175 0.013 0.081

Every pairwise gap here is under 1.5 SE over 160 episodes, so the ordering is not established โ€” the runs sit between 9 and 13 successes out of 160. avg reward is tau2's own all-or-nothing score: a run counts only if the database checks, action checks and assertions all pass.

Known caveat. One of the 40 test tasks appears in this training file. The synthesis step derives new tasks from training tasks by substituting faults, and in one case dropping a fault from a training task reproduced a test task's exact environment state (break_apn_settings + lock_sim_card_pin). The task ids still carry the training signature, so an id-based check does not catch it; the collision is only visible by comparing initial_state. The upper bound on the resulting inflation is +0.025 avg.

Usage

vllm serve MANGSEOK123/qwen3-4b-tau2-telecom-ours-1ep \
  --enable-auto-tool-choice --tool-call-parser hermes --max-model-len 40960

Evaluated at temperature=0, top_p=0.8, top_k=20 with tau2's vanilla LLMAgent and no experience in the prompt.

Downloads last month
21
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for MANGSEOK123/qwen3-4b-tau2-telecom-ours-1ep

Finetuned
(2375)
this model