topic-scope-v3

The model behind the topic_scope detector in flowx-border. Given a message and a taxonomy a policy supplies, it answers which node the message is about, or "none of these", with a calibrated probability.

The same design as flowxai/topic-scope-v2, on XLM-RoBERTa large instead of base: same corpus, same head, same calibration, and a training config that differs only in the learning rate and batch layout the larger encoder needs. Both are compared below on the same rows, with the bi-encoder both replace, flowxai/topic-scope.

How it runs

XLM-RoBERTa large, fine-tuned, plus a two-layer decision head trained from scratch. The late encoding: the message, the question and each node are encoded as separate short sequences, and meet only in the head. Node states are therefore computed once per taxonomy and cached; a scan encodes the message and runs the head.

file what bytes sha256
onnx/model.int8.onnx encoder, embedding Gather in int8, outputs last_hidden_state 1,466,403,115 483552daf200d566...
onnx/head.onnx decision head, fp32, takes hidden states and returns one logit per option 105,030,983 69a14a213934d71e...
decision_config.json question text, the none option, token budgets, none offset, temperatures 1,020 beaeab650cc02998...
tokenizer.json XLM-R tokenizer 17,098,086 acbd420e2269cdc1...
config.json encoder config 743 d44ba2614aaeeb01...

Token budgets: message 128, question 48, each node 64. A node is rendered as {key}: {text}, and the none option is always offered as none: none of these topics; the message is about something else. Its logit gets -0.5 before calibration. Validation fitted -0.25; -0.5 is shipped (changed 2026-10-01): topic-scope-v2 ships -0.5, and on 400 hand-written probes this model scores 0.752 at -0.5 against 0.728 at -0.25, while held-out accuracy moves by under 0.3 points (training repository, reports/topic_scope_v3_none_offset.json). Temperatures are as fitted. Probabilities are a softmax at a temperature fitted per option-count bucket on validation rows, listed in decision_config.json.

Options are packed without padding for the head: it has no position embedding, so this is the function it was trained as, over fewer positions.

Evaluation

Synthetic, and held out by deployment. The corpus is generated: 299 taxonomies and 135,747 messages in the 26 languages, written by claude-sonnet-5 (most messages) and gpt-5.6-luna, and every validation and test row checked by a judge model, a different model from the writer for 97% of messages. Splits are by taxonomy. unseen_type taxonomies come from deployment types no training taxonomy came from (airline, legal services, telecom); unseen_schema taxonomies are new taxonomies of trained types. Neither is a measurement on real traffic, and the numbers below should be read as how the model transfers to taxonomies it has not seen, not as a production error rate.

Accuracy is top-1 over the offered nodes plus none, after the none offset. The comparison is the model this one replaces, on the same rows, with its none bar (0.875) chosen on validation.

test this model topic-scope-v2 flowxai/topic-scope rows ECE
unseen_type 0.859 0.804 0.479 11,497 0.0118
unseen_schema 0.887 0.808 0.471 11,071 0.0141
v1 topic_scope test, recast 0.973 442 0.0145

Two training seeds. The published weights are seed 42. A second run on the same config and data, seed 1337, was evaluated on the same rows at the same none offset and not exported. The spread between them is the honest error bar on the table above.

test published seed seed 1337
topic_scope_v2 unseen_type 0.859 0.869
topic_scope_v2 unseen_schema 0.887 0.887
policy_questions unseen_type 0.807 0.823
policy_questions unseen_schema 0.817 0.802
policy minimal pairs, unseen_type 0.632 0.668

Hand-written probes

Every number above comes from synthetic rows written by the same models that wrote the training set. As a check from outside that style, 400 probes were written by hand: short, plain questions in 10 languages (en, ro, de, fr, es, it, pl, nl, pt, tr) against small taxonomies of 3 allowed and 3 disallowed nodes (banking, twice with two description styles, insurance and telecom), run through the library at threshold 0.5. in passes when nothing fires and the answer is not none; dis when the expected disallowed node fires; out, about nothing in the taxonomy, when nothing fires. The probes ship with the library as a test fixture.

model in dis out all
topic-scope-v2 127/200 76/120 80/80 0.708
topic-scope-v3 at -0.25 124/200 87/120 80/80 0.728
topic-scope-v3 at -0.5 130/200 91/120 80/80 0.752

Paired exact McNemar against topic-scope-v2: p=0.0474 over all 400 (46 probes only this model gets right, 28 only v2), p=0.0107 on the disallowed ones. The shipped offset was chosen after these probes were run, so they are not an independent test of that choice; the held-out synthetic rows, which moved by under 0.3 points, are.

Per language

Every language is reported, the weak ones included. Maltese is the weakest, as it is for every model on this base: Maltese is not in XLM-RoBERTa's pretraining set, and fine-tuning data cannot give the encoder a representation it never had. Irish is second.

language unseen_type topic-scope-v2 rows unseen_schema topic-scope-v2 rows
az 0.816 0.753 413 0.902 0.816 705
bg 0.825 0.804 445 0.899 0.812 485
cs 0.862 0.801 412 0.856 0.808 500
da 0.844 0.833 551 0.880 0.794 209
de 0.824 0.771 546 0.889 0.839 217
el 0.857 0.814 462 0.911 0.789 123
en 0.828 0.792 471 0.842 0.742 120
es 0.899 0.851 496 0.846 0.802 343
et 0.851 0.814 489 0.857 0.796 343
fi 0.885 0.851 495 0.885 0.798 573
fr 0.888 0.854 493 0.906 0.825 576
ga 0.808 0.710 427 0.848 0.676 732
hr 0.893 0.842 430 0.897 0.834 735
hu 0.859 0.830 383 0.895 0.843 477
it 0.895 0.847 372 0.905 0.839 465
lt 0.861 0.816 403 0.899 0.822 365
lv 0.895 0.805 389 0.887 0.797 364
mt 0.750 0.624 412 0.769 0.556 117
nl 0.862 0.818 406 0.839 0.750 112
pl 0.876 0.790 434 0.887 0.793 300
pt 0.895 0.851 449 0.874 0.824 301
ro 0.861 0.797 454 0.903 0.838 401
sk 0.896 0.793 425 0.929 0.837 406
sl 0.882 0.814 414 0.904 0.848 709
sv 0.888 0.823 419 0.896 0.840 686
tr 0.840 0.784 407 0.891 0.826 707

Per register

What kind of message it was. _hole rows had their correct node withheld from the options, so the right answer is none. negated_near_miss is a message that matches a node's words but falls under what the node excludes, and it is the weakest register on unseen types.

register unseen_type topic-scope-v2 rows unseen_schema topic-scope-v2 rows
in_domain_uncovered 0.883 0.865 814 0.912 0.872 693
in_node 0.869 0.806 4661 0.905 0.819 4485
in_node_hole 0.835 0.801 1521 0.848 0.766 1531
negated_near_miss 0.728 0.667 672 0.844 0.771 741
negated_near_miss_hole 0.759 0.782 87 0.905 0.778 63
out_of_taxonomy 0.980 0.979 817 0.997 0.995 645
sibling_near_miss 0.864 0.778 2212 0.886 0.795 2132
sibling_near_miss_hole 0.801 0.738 713 0.789 0.694 781

Export verification

300 test rows, the torch fp32 model on its training layout against the ONNX pipeline (onnxruntime CPU, one thread, int8 encoder, packed head): 0 changed their answer, and the largest logit difference was 0.13505. The head took 30.84 ms at p50 and 69.85 ms at p95 on those rows, one thread, on an Apple M-series laptop; the library's budget test measures the whole detector.

Latency

The whole topic_scope detector through the library, Apple M5 Max, one thread, onnxruntime CPU provider, measured 2026-09-30: tests/test_budgets.py p95 over 20 runs at REFERENCE_INPUT (396 characters), after one warm run; nodes are real descriptions from tests/fixtures/topic_scope/typed_26_languages.json, one disallowed and the rest allowed; node states cached per taxonomy as in a deployment. The shipped base model was measured beside it in the same session. Budget 300 ms.

nodes this model, p95 topic-scope-v2, p95
20 113.7 ms 57.4 ms
30 145.7 ms 82.7 ms
40 180.8 ms 102.3 ms

Base at 40 nodes reads 102.3 here against 104.1 recorded in the library's MEASURED_MS, so the machine was quiet.

Limits

  • Trained on taxonomies of 4 to 37 nodes. Head cost grows with the square of the total node text, so the library shortlists larger taxonomies to the nearest nodes by the model's own similarity term first, and records that it did.
  • A node description past 64 tokens is cut. Write the part that decides first.
  • The head was trained on four question families (topic_scope v1 and v2, injection and policy questions). The library asks it the topic question only. Its policy-question accuracy is not good enough to ship: 0.63 of minimal pairs on unseen types get both halves right, against a shipping bar of 0.80.
  • Short plain in-scope questions get "none of these" too often. On the hand-written probes, about 35% of questions about an allowed node do: 73 of 200 for topic-scope-v2, 70 of 200 for topic-scope-v3 at -0.5. None of them fires a disallowed node, so a policy that logs none, the library default, refuses nothing; under options.on_none: block each is a refusal. Neither the none offset nor a longer node description fixes it. Measure the rate on your own traffic before blocking on none.
  • Deterministic: no sampling. The same message, taxonomy and weights give the same answer.

Training

Source run: configs/typed_decisions_a3_large.yaml, base FacebookAI/xlm-roberta-large, 2 epochs, seed 42. Code and the plan it follows: docs/typed-decisions-plan.md in the training repository.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for flowxai/topic-scope-v3

Quantized
(15)
this model