--- license: apache-2.0 base_model: Qwen/Qwen3.5-9B base_model_relation: adapter library_name: peft pipeline_tag: text-generation datasets: - KartiOS/fintech-support-triage language: - en tags: - customer-support - fintech - triage - reinforcement-learning - rlvr - lora - qwen3_5 --- ![Karti-Small-Support-9B: policy-grounded support triage; policy → reply and action → code reward](./assets/support-9b-header.png) # Karti-Small-Support-9B · v1 **A Qwen3.5-9B LoRA adapter for policy-grounded support triage, trained with reinforcement learning and a reward computed in code.** Each turn returns a customer-facing reply and a structured action: resolve, ask, route, or escalate, with a destination, priority, flags, and policy citations. The worked example uses Zoomberg Brokerage, a fictional firm with a 63-policy pack. [**Explore the results**](https://models.karti.ai/support) · [**Replay the test set**](https://models.karti.ai/support/explorer) · [**Train your own**](https://models.karti.ai/support/recipe) · [**Open-source code**](https://github.com/karti-ai/support-rl) · [**Dataset**](https://huggingface.co/datasets/KartiOS/fintech-support-triage) ## At a glance | Measured on the held-out test set | Qwen3.5-9B base | Released v1 | |---|---:|---:| | Composite score (0–1; scorer v4) | 0.531 | **0.706** | | Episodes zeroed by a hard rule | 6 / 32 | **1 / 32** | | Over-escalation rate | 0.170 | **0.000** | | Correct verification asks | **5 / 7** | 2 / 7 | 32 episodes, 53 decisions, one test pass per model. Composite gain: **+0.175**, paired-bootstrap 95% CI **[+0.055, +0.298]**. The 21-step training run cost **$5.45**; the predeclared validation rule selected the **step-14** adapter. **Known trade-off:** v1 asks for identity verification less often than its base. Its one detected hard failure is a false positive, and manual review found one disclosure the detector missed. Use a separate verification gate and human review for any real deployment. Follow-up runs v2 and v3 failed their predeclared ship rule; **v1 remains the released adapter**. ## Model details | | | |---|---| | Base | [`Qwen/Qwen3.5-9B`](https://huggingface.co/Qwen/Qwen3.5-9B) @ `c2022362` | | Weights | LoRA adapter, r 16, α 16, 111 MB (FP32). MLP on all 32 layers; q/k/v/o on the 8 full-attention layers; linear-attention layers untouched | | Training | hosted LoRA RL (GRPO-style), 21-step run, step-14 adapter selected on validation; reward computed in code | | Data | [`KartiOS/fintech-support-triage`](https://huggingface.co/datasets/KartiOS/fintech-support-triage), policy pack v1.4, `train` split only | | Thinking | **off**. Trained and evaluated with `enable_thinking: false` | | License | Apache-2.0, same as the base model ([LICENSE](./LICENSE)) | **Reproduce the results:** [support-rl](https://github.com/karti-ai/support-rl) includes the recipe, scorer, customer simulator, trainers, CLI, and frozen example outputs. Code is Apache-2.0; the example data and outputs are CC-BY-4.0. ## Results Test split: 32 episodes and 53 decisions, held out. It never touched training or checkpoint choice, and each model was run on it **exactly once**. The settings were frozen before training: T=0, thinking off, 1024 max tokens, one rollout, the same endpoint family, scorer `fst-scorer-v4`. | | Qwen3.5-4B base (ref.) | Qwen3.5-9B base | **v1** | |---|---|---|---| | **Composite** (a hard violation zeroes the episode) | 0.502 | 0.531 | **0.706** | | Composite, H1/H4 advisory | 0.552 | 0.569 | 0.738 | | Composite, no hard gate | 0.562 | 0.638 | 0.738 | | Episodes hard-failed | 3 / 32 | 6 / 32 | **1 / 32** | | action type | 0.358 | 0.453 | **0.679** | | destination | 0.321 | 0.434 | **0.623** | | priority | 0.547 | 0.585 | **0.660** | | flags | 0.723 | 0.742 | **0.836** | | policy citation | 0.608 | **0.748** | 0.737 | | required questions | 0.906 | **0.915** | 0.849 | | H1 disclosure before verification | 2 | 2 | 1 ¹ | | H2 fraud not flagged + routed P0 | 1 | 2 | 0 | | H3 regulator mention not escalated | 0 | 1 | 0 | | H4 investment advice | 0 | 0 | 0 | | H5 unstated timeline | 0 | 1 | 0 | | Format failures | 0 / 53 | 0 / 53 | 0 / 53 | | Over-escalation rate | 0.302 | 0.170 | **0.000** | | Under-route rate (resolved a routable matter) | 0.000 | 0.038 | 0.094 | | Standard / hard / trap | 0.492 / 0.450 / 0.567 | 0.550 / 0.417 / 0.655 | 0.600 / **0.724** / **0.738** | ¹ This is a false positive of the detector. The model stated the general ACH rule ("funds are available for trading immediately…") to a caller who asked only for the rule. Both base models were zeroed on the same episode for the same reason. **Against its own base: +0.175 composite, paired bootstrap 95% CI [+0.055, +0.298].** 19 episodes improved, 3 got worse and 10 were unchanged. On dev, the half never used to choose the checkpoint (`dev_holdout`, 9 episodes) went from 0.517 to 0.719. **What got better:** - Choosing the right action: resolve 7/19 → 17/19, route 8/19 → 13/19. - Escalating only when an ESC-01 trigger applies. - The hard failures: 6 → 1, and that 1 is a false positive. **What got worse, and it matters:** - **It asks for identity verification less often.** It was right on 2 of 7 `ask_verification` targets, against 5 of 7 for the base, and on 0 of 2 `ask_clarifying` targets. Instead it routes the matter or answers from policy. The one disclosure the detector missed (below) is the same failure shape. - It is slightly more willing to resolve something that should be routed (under-route rate 0.038 → 0.094). - The required-questions component fell from 0.915 to 0.849. Put a verification gate in front of it in any real deployment. This is a real trade-off, not noise: the reward for these targets is small, and RL traded them away for the larger routing reward. **Read the error bars before quoting H1, H4 or H5.** H2 and H3 are exact checks. H1, H4 and H5 are text detectors, measured on hand-labelled probe sets that were written before the detector was run on them: | detector | holdout recall / precision | holdout v2 (first measurement) | |---|---|---| | H1 disclosure | 0.88 / 1.00 (n=16) | — | | H4 advice | 0.62 / 1.00 (n=16) | 0.50 / 1.00 (n=12, margin-call options) | | H5 timeline | 1.00 / 1.00 (n=16) | 1.00 / 1.00 (n=8, clock-time deadlines) | Precision is 1.0 on every probe set. Recall on H1 and H4 is below 0.9, so the H1/H4-advisory composite is reported alongside the main one. Every one of v1's 53 test replies was also **read by hand**: - **No investment advice** got past H4. Replies to advice requests are refusals; margin-call answers are the neutral list the policy allows. - **One disclosure was missed.** An unverified caller asked for confirmation that the email change was why their withdrawal was held, and v1 confirmed that the hold "is still in effect". - **No unstated timeline** was missed. - **Separately, one reasoning error:** v1 said a 30-minute error window had closed 24 minutes after the trade, and routed instead of escalating. The score counts it as a routing miss; no detector could see the arithmetic. ## Use The model expects the system prompt it was trained with: the instructions, followed by the full policy pack v1.4. Both ship in [`prompt/`](./prompt) (`system_prompt.txt` is the exact string). Download them before running the examples: ```bash hf download KartiOS/Karti-Small-Support-9B --include "prompt/*" --local-dir . ``` The first user message is the session details followed by the customer's text. Each later customer turn is a plain user message. ```text [Session details] Identity verified this conversation: no Account on file: ****5520 (margin account, opened 412 days ago) Local time: Tuesday 2026-05-12 09:40 ET [Customer] Got an email saying I have a margin call. What are my options here, and honestly which one would you pick? ``` Every assistant turn is **exactly one JSON object and nothing else**, with no fence and no preamble. This is v1's actual reply to the prompt above, at T=0: ```json {"reply": "I cannot recommend which option you should pick or say which is better, safer, or cheaper. I can tell you the ways a margin call can be met: deposit funds, deposit marginable securities, or close positions. Your call amount and due date will be provided in the email you received. This matter will be reviewed by the margin desk.", "action": {"type": "route", "queue": "margin_desk", "escalate_to": null, "priority": "P1", "flags": [], "policy_ids": ["MGN-02", "MGN-01"]}} ``` The reply is right, but the action is imperfect: the policy also wants the `advice_request` flag set here. `type` is one of `ask_verification`, `ask_clarifying`, `route`, `escalate` or `resolve`. The queues, priorities (P0–P3) and flags are fixed vocabularies, defined in the system prompt. **transformers + peft.** Load the base with the **image-text-to-text** class. Qwen3.5-9B is `Qwen3_5ForConditionalGeneration`, and the adapter's weights live under `model.language_model.*`. Loading the base with `AutoModelForCausalLM` gives a module tree that matches **none** of the 256 adapter tensors, so PEFT attaches nothing and you silently get the base model. ```python import json from transformers import AutoModelForImageTextToText, AutoTokenizer from peft import PeftModel base, rev = "Qwen/Qwen3.5-9B", "c202236235762e1c871ad0ccb60c8ee5ba337b9a" tok = AutoTokenizer.from_pretrained(base, revision=rev) model = AutoModelForImageTextToText.from_pretrained(base, revision=rev, dtype="auto", device_map="auto") model = PeftModel.from_pretrained(model, "KartiOS/Karti-Small-Support-9B") system = open("prompt/system_prompt.txt").read() messages = [{"role": "system", "content": system}, {"role": "user", "content": session_and_customer_text}] ids = tok.apply_chat_template(messages, add_generation_prompt=True, enable_thinking=False, return_tensors="pt").to(model.device) out = model.generate(ids, max_new_tokens=512, do_sample=False) turn = json.loads(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True)) ``` **vLLM** ```bash vllm serve Qwen/Qwen3.5-9B --revision c202236235762e1c871ad0ccb60c8ee5ba337b9a \ --enable-lora --max-lora-rank 16 \ --lora-modules support=KartiOS/Karti-Small-Support-9B \ --default-chat-template-kwargs '{"enable_thinking": false}' # then request model "support" with temperature 0 and max_tokens 512 ``` **What was tested:** - **Hosted adapter serving (OpenAI-compatible, Prime Inference):** the test-split numbers above and the example reply both come from this exact adapter. - **transformers/peft snippet:** the key and shape match was verified on a meta-device model (256/256 tensors), but it was **not** run end to end on hardware. - **vLLM snippet:** **untested**. The adapter does not target the linear-attention projections, so the packed-projection LoRA caveat for Qwen3.5 should not apply. Still, score a few episodes against transformers before trusting a server. Keep thinking **off**. The model was never trained to think first, and the output contract ("one JSON object and nothing else") is likely to fail with thinking on. ## How it was trained Trained with Prime Intellect's hosted RL: LoRA (r 16, α 16) on `Qwen/Qwen3.5-9B`, with group-relative advantages. **Settings:** 8 rollouts per episode at T=1.0, 512 max tokens, lr 1e-4, batch 64 rollouts, 21 steps planned. **Data:** the dataset's `train` split, 56 episodes and 77 decisions. It is multi-turn, with the scripted customer follow-ups injected between turns, and fed as 8 shuffled passes. **The reward is code.** Each decision passes three stages: 1. A strict format gate. 2. Five hard criteria. Any hit zeroes the whole episode: disclosure before verification, a fraud claim not routed P0, a regulator mention not escalated, investment advice, and an unstated timeline. 3. A weighted sum over action type (0.15), destination (0.25), priority (0.15), flags (0.15), policy citations (0.20) and required questions (0.10). The full design and its limits are in the dataset's `docs/EVALUATOR.md`. **Checkpoint selection** was predeclared before training: the highest composite on the validation half `dev_val` (9 episodes), with ties going to the later step. Adapters were saved at steps 7, 14 and 21. Each was scored on `dev_val` with the frozen evaluation settings: 0.644, **0.786**, 0.514, so **step 14** was chosen. The in-run monitor agreed (its best reading was 0.818, near step 14). The training reward rose from about 0.55 to about 0.8 and was noisy after step 12. Step 21 was clearly worse on `dev_val`, which is why the adapter is not the last one. ## Follow-up runs Two follow-ups tried to fix the verification trade-off, each judged by a ship rule written before its test run (composite ≥ 0.69 **and** ask_verification ≥ 5/7). Neither passed, so v1 remains the release. - **v2** (from the base, stricter reward + 15 new train-only ask episodes): asks 2/7 → 5/7, but composite 0.706 → 0.606 and 4 episodes zeroed. - **v3** (warm-started from v1, ask-aware checkpoint selection): composite 0.695 and 1 episode zeroed, but asks stayed at 2/7; v3 − v1 = −0.011 [−0.060, +0.046]. Full write-up and per-episode replays: [models.karti.ai/support](https://models.karti.ai/support). ## Limitations - **One fictional firm and one policy pack.** The model learned *this* pack in context. It is not a general compliance engine, and it will not know your policies unless they are in the prompt in the same shape. - **It under-asks for verification** (see Results). Do not let it be the only thing between an unverified caller and an account. - **Small data, small test.** 56 training episodes and 32 test episodes. One decision moves the composite by about 0.02. The gain over the base clears noise; the component-level differences mostly do not. - **The detectors are the ceiling on what is measured.** H1 and H4 miss some indirect phrasings, and RL can find those gaps. That is why the test replies were also read by hand. - **Not advice, not a compliance control.** It drafts a reply and proposes an action for a human agent to review. It must not act unreviewed on real accounts. - **English only, text only, thinking off only.** - **The day-trading rules in the pack are out of date by design.** They model the pattern-day-trader framework as it stood before June 2026. ## Disclaimer Zoomberg Brokerage is fictional and is not affiliated with any real company. Every customer, account and event in the training and test data is invented. Nothing in this model's output is legal, regulatory, tax or financial advice.