Qwen3-4B-Instruct-2507 โ tau2-bench telecom, externally authored memories ("ours")
Qwen3-4B-Instruct-2507 consolidated with OEL on the tau2-bench telecom domain, using memories written by a separate authoring pipeline rather than extracted from the model's own rollouts. The companion model qwen3-4b-tau2-telecom-oel-1ep is the self-extracted counterpart, trained the same way.
How it was made
- Deploy โ the student plays each of the 100 taskโmemory pairs without memory; every agent turn is saved.
- Consolidate โ the same trajectories are scored again with each task's memory injected into the system prompt (the teacher), and the memory-free student is fit to that distribution with a token-level full KL loss. Teacher and student share weights; the only difference is the prompt. No reward signal is used.
| base model | Qwen3-4B-Instruct-2507 |
| training pairs | 100 (telecom_final_pairs_format.json, 50 source tasks ร 2 memories) |
| batch / steps | 12 / 9 (1 epoch) |
| learning rate | 3e-6, constant |
| loss | full KL, all response tokens |
| user simulator | gpt-4.1-mini (temperature 0) |
Results โ tau2-bench telecom test (40 tasks, 4 trials, no memory at eval)
| pass@3 | pass@4 | pass^3 | avg reward | |
|---|---|---|---|---|
| base Qwen3-4B-Instruct-2507 | 0.144 | 0.175 | 0.000 | 0.056 |
| this model (external memories) | 0.144 | 0.175 | 0.006 | 0.062 |
| self-extracted, 1 epoch | 0.156 | 0.175 | 0.013 | 0.081 |
Every pairwise gap here is under 1.5 SE over 160 episodes, so the ordering is not
established โ the runs sit between 9 and 13 successes out of 160. avg reward is tau2's
own all-or-nothing score: a run counts only if the database checks, action checks and
assertions all pass.
Known caveat. One of the 40 test tasks appears in this training file. The synthesis
step derives new tasks from training tasks by substituting faults, and in one case
dropping a fault from a training task reproduced a test task's exact environment state
(break_apn_settings + lock_sim_card_pin). The task ids still carry the training
signature, so an id-based check does not catch it; the collision is only visible by
comparing initial_state. The upper bound on the resulting inflation is +0.025 avg.
Usage
vllm serve MANGSEOK123/qwen3-4b-tau2-telecom-ours-1ep \
--enable-auto-tool-choice --tool-call-parser hermes --max-model-len 40960
Evaluated at temperature=0, top_p=0.8, top_k=20 with tau2's vanilla LLMAgent
and no experience in the prompt.
- Downloads last month
- 21
Model tree for MANGSEOK123/qwen3-4b-tau2-telecom-ours-1ep
Base model
Qwen/Qwen3-4B-Instruct-2507