--- license: apache-2.0 base_model: Qwen/Qwen3-4B-Instruct-2507 tags: - tau2-bench - telecom - oel - experience-distillation --- # Qwen3-4B-Instruct-2507 — tau2-bench telecom, externally authored memories ("ours") Qwen3-4B-Instruct-2507 consolidated with OEL on the tau2-bench **telecom** domain, using memories written by a separate authoring pipeline rather than extracted from the model's own rollouts. The companion model [qwen3-4b-tau2-telecom-oel-1ep](https://huggingface.co/MANGSEOK123/qwen3-4b-tau2-telecom-oel-1ep) is the self-extracted counterpart, trained the same way. ## How it was made 1. **Deploy** — the student plays each of the 100 task–memory pairs *without* memory; every agent turn is saved. 2. **Consolidate** — the same trajectories are scored again with each task's memory injected into the system prompt (the teacher), and the memory-free student is fit to that distribution with a token-level full KL loss. Teacher and student share weights; the only difference is the prompt. No reward signal is used. | | | |---|---| | base model | Qwen3-4B-Instruct-2507 | | training pairs | 100 (`telecom_final_pairs_format.json`, 50 source tasks × 2 memories) | | batch / steps | 12 / 9 (1 epoch) | | learning rate | 3e-6, constant | | loss | full KL, all response tokens | | user simulator | gpt-4.1-mini (temperature 0) | ## Results — tau2-bench telecom test (40 tasks, 4 trials, no memory at eval) | | pass@3 | pass@4 | pass^3 | avg reward | |---|---|---|---|---| | base Qwen3-4B-Instruct-2507 | 0.144 | 0.175 | 0.000 | 0.056 | | **this model** (external memories) | 0.144 | 0.175 | 0.006 | 0.062 | | self-extracted, 1 epoch | 0.156 | 0.175 | 0.013 | 0.081 | Every pairwise gap here is under 1.5 SE over 160 episodes, so the ordering is not established — the runs sit between 9 and 13 successes out of 160. `avg reward` is tau2's own all-or-nothing score: a run counts only if the database checks, action checks and assertions all pass. **Known caveat.** One of the 40 test tasks appears in this training file. The synthesis step derives new tasks from training tasks by substituting faults, and in one case dropping a fault from a training task reproduced a test task's exact environment state (`break_apn_settings` + `lock_sim_card_pin`). The task ids still carry the training signature, so an id-based check does not catch it; the collision is only visible by comparing `initial_state`. The upper bound on the resulting inflation is +0.025 avg. ## Usage ```bash vllm serve MANGSEOK123/qwen3-4b-tau2-telecom-ours-1ep \ --enable-auto-tool-choice --tool-call-parser hermes --max-model-len 40960 ``` Evaluated at `temperature=0, top_p=0.8, top_k=20` with tau2's vanilla `LLMAgent` and no experience in the prompt.