Instructions to use monte-inc/qwen3.5-35b-a3b-banking-sft-v6 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use monte-inc/qwen3.5-35b-a3b-banking-sft-v6 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Qwen3.5-35B-A3B") model = PeftModel.from_pretrained(base_model, "monte-inc/qwen3.5-35b-a3b-banking-sft-v6") - Notebooks
- Google Colab
- Kaggle
qwen3.5-35b-a3b-banking-sft-v6
QLoRA adapter for unsloth/Qwen3.5-35B-A3B trained on the tau2-bench banking_knowledge v2 data (Claude Opus/Sonnet + GPT-5.2 trajectories with reasoning traces).
This is the first SFT run that does not degrade tool-calling on the tau2-bench banking domain — earlier attempts (v1–v5) all produced 0% pass rate due to masking/data/rank issues. See docs/sft-history.md for the post-mortem.
Training
- Framework: Unsloth + TRL 0.22.2 (Transformers 5.2.0, PyTorch 2.9.0 + CUDA 12.8)
- Config:
training/run_qwen_sft.pyinMonte-Inc/tau2-banking-sft@main - Hardware: GH200 480 GB (aarch64, vLLM 0.19.0)
- Runtime: 1 epoch, 622 steps, ~61 min (5.5 s/step —
torch._grouped_mmMoE path in vanilla torch 2.9 cut training ~10× vs v5) - LoRA: rank 32, alpha 64, dropout 0
- Targets: attention
{q,k,v,o}_proj+ MLP{gate,up,down,gate_up}_proj(regular dense LoRA) plus MoE expert 3D fused params (mlp.experts.gate_up_proj,mlp.experts.down_proj) — auto-added by Unsloth's PEFT path for Qwen3.5 MoE - Hyperparams: lr 2e-4, cosine schedule, 62 warmup steps, optim
adamw_8bit, weight_decay 0.001, seq_len 8192, batch 1 × grad_accum 4 (effective 4), packing off - Quantization: 4-bit base (bnb double-quant) + bf16 mixed precision
- Final training loss: 0.048 (step 622); first step 0.67
Repo contents
adapter_model.safetensors(7.4 GB) — PEFT LoRA weights (includes MoE expert deltas)adapter_config.json—target_parameters: [mlp.experts.gate_up_proj, mlp.experts.down_proj],r=32, alpha=64chat_template.jinja— Qwen 3.5 native template (no patches needed; works as-is for tau2-bench/litellm tool-call rendering)tokenizer.json,tokenizer_config.json— Qwen 3.5 tokenizer (kept self-contained for downstream loading)trainer_state.json— full step-by-step loss + grad-norm + LR historytraining_args.json— exactSFTConfigused (verbatim dump oftrainer.args.to_dict())
Inference setup
Merge the adapter (Unsloth → vLLM)
python -c "
from unsloth import FastLanguageModel
from peft import PeftModel
model, tok = FastLanguageModel.from_pretrained(
'unsloth/Qwen3.5-35B-A3B', max_seq_length=8192, load_in_4bit=False,
)
model = PeftModel.from_pretrained(model, 'monte-inc/qwen3.5-35b-a3b-banking-sft-v6')
model = model.merge_and_unload()
model.save_pretrained_merged('./merged-qwen-sft-v6', tok, save_method='merged_16bit')
"
Output is ~66 GB (bf16).
Serve with vLLM
vllm serve ./merged-qwen-sft-v6 \
--served-model-name qwen-banking-sft \
--host 0.0.0.0 --port 8000 \
--dtype bfloat16 \
--max-model-len 100000 \
--gpu-memory-utilization 0.90 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--enforce-eager \
--num-gpu-blocks-override 256 \
--limit-mm-per-prompt '{"image":0,"video":0}'
Known issues
1. torch.compile crashes on Grace Hopper MoE
Symptom: vLLM's MoE compiled path hits pytorch/pytorch#176178 on GH200, also hangs KV-cache profiling. Workaround: --enforce-eager --num-gpu-blocks-override 256 (in the vLLM command above). Eval is ~3× slower in eager mode but stable.
2. --limit-mm-per-prompt required for Qwen3.5 MoE
Without it, vLLM provisions multimodal heads that this text-only adapter doesn't use, wasting ~8 GB of KV cache. The '{"image":0,"video":0}' flag disables those heads.
Baselines on banking_knowledge (gpt-5.2 user sim, terminal_use, 4 trials × 200 max steps × seed 42)
| Run | Pass rate |
|---|---|
monte-inc/qwen3.5-35b-a3b-banking-sft-v6 (this model) |
13/94 (13.8%) — 1 trial, 95/97 tasks ran |
unsloth/gemma-4-26B-A4B-it (base, thinking on) |
53/382 (13.9%) — 4 trials |
monte-inc/gemma4-26b-a4b-banking-sft-v3 |
(see that repo) |
Qwen/Qwen3.5-35B-A3B (base, thinking on) |
20/322 (6.2%) — gpt-4.1 user sim |
This SFT run roughly matches the Gemma 4 26B MoE baseline despite Qwen3.5 base being significantly weaker (~6.2% with gpt-4.1 user sim) — the v2 data (reasoning extraction + Claude-only) plus rank 32 closed the gap.
tau2-bench eval command
tau2 run --domain banking_knowledge \
--agent-llm openai/qwen-banking-sft \
--agent-llm-args '{"api_base":"http://localhost:8000/v1","temperature":0.0}' \
--user-llm gpt-5.2 \
--retrieval-config terminal_use \
--max-concurrency 10 --num-trials 4 --max-steps 200 \
--seed 42 --max-retries 3 \
--save-to qwen-sft-v6-banking
Source code
Trained from Monte-Inc/tau2-banking-sft@main on branch main (commit ced327c at time of publish). The exact training script + data pipeline:
- Training:
training/run_qwen_sft.py - Data:
data/banking/v2/banking_sft_ready.jsonl - Pipeline:
DATA_PIPELINE.md - Eval results:
benchmarks/results/sft/qwen-sft-v6-banking/results.json
For the post-mortem of failed earlier runs (v1–v5), see docs/sft-history.md.
- Downloads last month
- 10
Model tree for monte-inc/qwen3.5-35b-a3b-banking-sft-v6
Base model
Qwen/Qwen3.5-35B-A3B-Base