qwen3.5-35b-a3b-banking-sft-v6

QLoRA adapter for unsloth/Qwen3.5-35B-A3B trained on the tau2-bench banking_knowledge v2 data (Claude Opus/Sonnet + GPT-5.2 trajectories with reasoning traces).

This is the first SFT run that does not degrade tool-calling on the tau2-bench banking domain — earlier attempts (v1–v5) all produced 0% pass rate due to masking/data/rank issues. See docs/sft-history.md for the post-mortem.

Training

  • Framework: Unsloth + TRL 0.22.2 (Transformers 5.2.0, PyTorch 2.9.0 + CUDA 12.8)
  • Config: training/run_qwen_sft.py in Monte-Inc/tau2-banking-sft@main
  • Hardware: GH200 480 GB (aarch64, vLLM 0.19.0)
  • Runtime: 1 epoch, 622 steps, ~61 min (5.5 s/step — torch._grouped_mm MoE path in vanilla torch 2.9 cut training ~10× vs v5)
  • LoRA: rank 32, alpha 64, dropout 0
  • Targets: attention {q,k,v,o}_proj + MLP {gate,up,down,gate_up}_proj (regular dense LoRA) plus MoE expert 3D fused params (mlp.experts.gate_up_proj, mlp.experts.down_proj) — auto-added by Unsloth's PEFT path for Qwen3.5 MoE
  • Hyperparams: lr 2e-4, cosine schedule, 62 warmup steps, optim adamw_8bit, weight_decay 0.001, seq_len 8192, batch 1 × grad_accum 4 (effective 4), packing off
  • Quantization: 4-bit base (bnb double-quant) + bf16 mixed precision
  • Final training loss: 0.048 (step 622); first step 0.67

Repo contents

  • adapter_model.safetensors (7.4 GB) — PEFT LoRA weights (includes MoE expert deltas)
  • adapter_config.json — target_parameters: [mlp.experts.gate_up_proj, mlp.experts.down_proj], r=32, alpha=64
  • chat_template.jinja — Qwen 3.5 native template (no patches needed; works as-is for tau2-bench/litellm tool-call rendering)
  • tokenizer.json, tokenizer_config.json — Qwen 3.5 tokenizer (kept self-contained for downstream loading)
  • trainer_state.json — full step-by-step loss + grad-norm + LR history
  • training_args.json — exact SFTConfig used (verbatim dump of trainer.args.to_dict())

Inference setup

Merge the adapter (Unsloth → vLLM)

python -c "
from unsloth import FastLanguageModel
from peft import PeftModel
model, tok = FastLanguageModel.from_pretrained(
    'unsloth/Qwen3.5-35B-A3B', max_seq_length=8192, load_in_4bit=False,
)
model = PeftModel.from_pretrained(model, 'monte-inc/qwen3.5-35b-a3b-banking-sft-v6')
model = model.merge_and_unload()
model.save_pretrained_merged('./merged-qwen-sft-v6', tok, save_method='merged_16bit')
"

Output is ~66 GB (bf16).

Serve with vLLM

vllm serve ./merged-qwen-sft-v6 \
  --served-model-name qwen-banking-sft \
  --host 0.0.0.0 --port 8000 \
  --dtype bfloat16 \
  --max-model-len 100000 \
  --gpu-memory-utilization 0.90 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --enforce-eager \
  --num-gpu-blocks-override 256 \
  --limit-mm-per-prompt '{"image":0,"video":0}'

Known issues

1. torch.compile crashes on Grace Hopper MoE

Symptom: vLLM's MoE compiled path hits pytorch/pytorch#176178 on GH200, also hangs KV-cache profiling. Workaround: --enforce-eager --num-gpu-blocks-override 256 (in the vLLM command above). Eval is ~3× slower in eager mode but stable.

2. --limit-mm-per-prompt required for Qwen3.5 MoE

Without it, vLLM provisions multimodal heads that this text-only adapter doesn't use, wasting ~8 GB of KV cache. The '{"image":0,"video":0}' flag disables those heads.

Baselines on banking_knowledge (gpt-5.2 user sim, terminal_use, 4 trials × 200 max steps × seed 42)

Run Pass rate
monte-inc/qwen3.5-35b-a3b-banking-sft-v6 (this model) 13/94 (13.8%) — 1 trial, 95/97 tasks ran
unsloth/gemma-4-26B-A4B-it (base, thinking on) 53/382 (13.9%) — 4 trials
monte-inc/gemma4-26b-a4b-banking-sft-v3 (see that repo)
Qwen/Qwen3.5-35B-A3B (base, thinking on) 20/322 (6.2%) — gpt-4.1 user sim

This SFT run roughly matches the Gemma 4 26B MoE baseline despite Qwen3.5 base being significantly weaker (~6.2% with gpt-4.1 user sim) — the v2 data (reasoning extraction + Claude-only) plus rank 32 closed the gap.

tau2-bench eval command

tau2 run --domain banking_knowledge \
  --agent-llm openai/qwen-banking-sft \
  --agent-llm-args '{"api_base":"http://localhost:8000/v1","temperature":0.0}' \
  --user-llm gpt-5.2 \
  --retrieval-config terminal_use \
  --max-concurrency 10 --num-trials 4 --max-steps 200 \
  --seed 42 --max-retries 3 \
  --save-to qwen-sft-v6-banking

Source code

Trained from Monte-Inc/tau2-banking-sft@main on branch main (commit ced327c at time of publish). The exact training script + data pipeline:

For the post-mortem of failed earlier runs (v1–v5), see docs/sft-history.md.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for monte-inc/qwen3.5-35b-a3b-banking-sft-v6

Adapter
(12)
this model

Collection including monte-inc/qwen3.5-35b-a3b-banking-sft-v6