Token Classification
Transformers
ONNX
Safetensors
modernbert
ner
on-device
privacy
flowx
openner
banking
de-identification
Instructions to use flowxai/ibandetect with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use flowxai/ibandetect with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="flowxai/ibandetect")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("flowxai/ibandetect") model = AutoModelForTokenClassification.from_pretrained("flowxai/ibandetect", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Remove a false checksum guarantee, say what the 1.0000 F1 is, state the training domain
Browse filesThe card claimed identifiers are validated by checksum, as 'a correctness guarantee general LLMs lack'. No repository in the OpenNER family ships validation code, and a ModernBERT encoder cannot run a check digit, so the claim was false for every model in the family.
The held-out F1 is now explained rather than headlined: train and test come from one generator, and two models from the same pipeline (filingtag 0.6444, ibandetect 0.9488) show it can report below 1.0, so a 1.0 means the generated task was trivially separable.
Also removes a pointer to an OpenNER benchmark that is not published, and adds a 'Trained on' line naming the actual domain.
README.md
CHANGED
|
@@ -24,11 +24,26 @@ metrics:
|
|
| 24 |
- **Task:** token-classification
|
| 25 |
- **Base model:** `answerdotai/ModernBERT-base`
|
| 26 |
- **Entity types (4):** ACCOUNT, IBAN, ROUTING, SWIFT
|
| 27 |
-
- **
|
| 28 |
- **Runtime:** CPU, Apple Silicon, one GPU, or browser/edge via ONNX (INT8). ~100-160 ms/doc on CPU.
|
| 29 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
## Why a small model
|
| 31 |
-
Fine-tuned encoders match or beat frontier LLMs on structured, convention-bound extraction, at a fraction of the latency and cost, with **zero data egress**.
|
| 32 |
|
| 33 |
## Usage
|
| 34 |
```python
|
|
@@ -40,4 +55,4 @@ model = AutoModelForTokenClassification.from_pretrained("flowxai/ibandetect")
|
|
| 40 |
## License & attribution
|
| 41 |
Licensed under the **Apache License 2.0**. Copyright 2026 **FlowX.AI** (https://flowx.ai). See the `NOTICE` file. Trained on synthetic, checksum-validated data.
|
| 42 |
|
| 43 |
-
_Part of the FlowX OpenNER model family.
|
|
|
|
| 24 |
- **Task:** token-classification
|
| 25 |
- **Base model:** `answerdotai/ModernBERT-base`
|
| 26 |
- **Entity types (4):** ACCOUNT, IBAN, ROUTING, SWIFT
|
| 27 |
+
- **Trained on:** synthetic banking payment records.
|
| 28 |
- **Runtime:** CPU, Apple Silicon, one GPU, or browser/edge via ONNX (INT8). ~100-160 ms/doc on CPU.
|
| 29 |
|
| 30 |
+
## Evaluation
|
| 31 |
+
|
| 32 |
+
**Held-out F1 on synthetic data: 0.9488.** Train and test are drawn from the same generator, so this describes performance on that generator's distribution rather than on your documents.
|
| 33 |
+
|
| 34 |
+
It is one of only two models in this family not reporting exactly 1.0000, which is worth reading in both directions. It is a weaker number than its siblings publish, and a more informative one: a 1.0 elsewhere in this family reflects a generated task that was trivially separable rather than a better model.
|
| 35 |
+
|
| 36 |
+
There is no per-label breakdown, only this aggregate, so it cannot tell you which labels carry the score. **Validate on your own documents before production use.**
|
| 37 |
+
|
| 38 |
+
## What this model does not do
|
| 39 |
+
It labels spans. It does not validate them, and nothing in this repository does: no check digit is verified anywhere here, not IBAN mod-97, not the Luhn algorithm, not ISIN, LEI, VIN or ISO-6346. A span this model labels as an identifier may not be a valid identifier. Validate downstream if your use needs that guarantee.
|
| 40 |
+
|
| 41 |
+
**Correction, 2026-09-14.** Until this date the card claimed that identifiers are "validated by checksum (IBAN mod-97, card Luhn, ISIN/LEI, container ISO-6346, VIN, national IDs), a correctness guarantee general LLMs lack". That sentence was shared boilerplate across the OpenNER family and it was not true of any model in it. These repositories contain a config, weights, an ONNX export, a tokenizer and a metrics file, and no validation code of any kind. If you relied on that sentence, the guarantee it described does not exist and never did.
|
| 42 |
+
|
| 43 |
+
The licence note at the foot of this card says the model was trained on "synthetic, checksum-validated data". That is a statement about how the **training corpus** was generated. It is not a statement about anything this model checks when you run it, and the two were being read as one claim.
|
| 44 |
+
|
| 45 |
## Why a small model
|
| 46 |
+
Fine-tuned encoders match or beat frontier LLMs on structured, convention-bound extraction, at a fraction of the latency and cost, with **zero data egress**.
|
| 47 |
|
| 48 |
## Usage
|
| 49 |
```python
|
|
|
|
| 55 |
## License & attribution
|
| 56 |
Licensed under the **Apache License 2.0**. Copyright 2026 **FlowX.AI** (https://flowx.ai). See the `NOTICE` file. Trained on synthetic, checksum-validated data.
|
| 57 |
|
| 58 |
+
_Part of the FlowX OpenNER model family._
|