MaziyarPanahi commited on
Commit
5515f8a
·
verified ·
1 Parent(s): 95799e5

Add Ministral-3B-PII: GRPO-trained structured PII extraction (text-to-text)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,1591 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: transformers
3
+ license: apache-2.0
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - pii
7
+ - ner
8
+ - privacy
9
+ - compliance
10
+ - hipaa
11
+ - gdpr
12
+ - pci-dss
13
+ - multilingual
14
+ - structured-output
15
+ - grpo
16
+ language:
17
+ - en
18
+ - zh
19
+ - hi
20
+ - es
21
+ - ar
22
+ - fr
23
+ - bn
24
+ - ru
25
+ - pt
26
+ - ja
27
+ - de
28
+ - ko
29
+ - it
30
+ - tr
31
+ - vi
32
+ - fa
33
+ - pl
34
+ - nl
35
+ - sw
36
+ - th
37
+ ---
38
+
39
+ # Ministral-3B-PII
40
+
41
+ **Ministral-3B-PII** is a 3.3B-parameter language model that detects **personally identifiable information (PII)** in unstructured text and returns it as a **structured JSON array** of typed entities. Give it any text and it emits a list of `{"text": ..., "label": ...}` objects spanning 69 PII entity types across the healthcare, financial, identity, and digital domains.
42
+
43
+ The model is an **experimental, reinforcement-learning–trained** variant of a Ministral-3B base. It was optimized with **GRPO** (Group Relative Policy Optimization) specifically to produce valid, schema-consistent JSON and to detect PII with high precision — making it suited to redaction, de-identification, and compliance workflows (HIPAA, GDPR, PCI-DSS).
44
+
45
+ > **Research preview.** This is an experimental model intended for evaluation and pipeline integration. Use it as one layer in a broader privacy/compliance system, not as a sole compliance control.
46
+
47
+ > **⚠️ Text input only.** This release is a **text-to-text** model: it reads text and returns JSON. The underlying architecture also contains a vision encoder, but **image-to-text PII extraction is not supported** in this version — passing images is not a validated path. Multimodal (image → PII) support is planned for a future release.
48
+
49
+ ## Key Results
50
+
51
+ Evaluated on a 1,000-sample held-out PII benchmark with greedy decoding, a 2,048-token prompt budget, and no assistant-side JSON-fence prefill.
52
+
53
+ | Metric | Score |
54
+ |---|---|
55
+ | Valid JSON rate | 1.000 |
56
+ | Valid label rate | 0.975 |
57
+ | Micro precision | 0.914 |
58
+ | Micro recall | 0.859 |
59
+ | **Micro F1** | **0.886** |
60
+ | Format consistency | 100% |
61
+ | Empty-output consistency | 100% |
62
+
63
+ Every generation parsed as valid JSON, and the model reliably returns `[]` for text containing no PII.
64
+
65
+ ## Supported PII Labels
66
+
67
+ The model recognizes **69 PII entity types**. Each detected span is returned as `{"text": "...", "label": "..."}` using the label names below.
68
+
69
+ <details>
70
+ <summary><b>View all 69 entity types by category</b></summary>
71
+
72
+ <br>
73
+
74
+ | Category | Entity types |
75
+ |---|---|
76
+ | **Identity &amp; demographics** | `first_name`, `last_name`, `title`, `date_of_birth`, `age`, `gender`, `nationality`, `race`, `ethnicity`, `race_ethnicity`, `religion`, `religious_belief`, `marital_status`, `sexuality`, `political_view`, `language`, `biometric_identifier` |
77
+ | **Contact &amp; address** | `email`, `phone_number`, `fax_number`, `street_address`, `building_number`, `city`, `county`, `state`, `postcode`, `zip_code`, `country`, `coordinate` |
78
+ | **Government &amp; legal IDs** | `social_security_number`, `ssn`, `national_id`, `driver_license_number`, `tax_id`, `license_plate`, `vehicle_identifier`, `certificate_license_number`, `unique_id` |
79
+ | **Healthcare** | `medical_record_number`, `health_plan_beneficiary_number`, `blood_type` |
80
+ | **Financial** | `credit_debit_card`, `cvv`, `pin`, `account_number`, `bank_routing_number`, `iban`, `swift_bic`, `salary` |
81
+ | **Employment &amp; organization** | `occupation`, `employment_status`, `employee_id`, `education_level`, `organization`, `company_name`, `customer_id` |
82
+ | **Digital &amp; network** | `ip_address`, `ipv4`, `ipv6`, `mac_address`, `url`, `user_name`, `password`, `http_cookie`, `api_key`, `device_identifier` |
83
+ | **Temporal** | `date`, `date_time`, `time` |
84
+
85
+ </details>
86
+
87
+ ## Quickstart
88
+
89
+ The PII extraction system prompt (with few-shot examples) is baked into the chat template, so no system message is required — just send the text. The template does **not** prefill a markdown `json` fence; the model emits the JSON array itself.
90
+
91
+ ```python
92
+ import torch
93
+ from transformers import AutoModelForImageTextToText, AutoTokenizer
94
+
95
+ model_id = "OpenMed/Ministral-3B-PII"
96
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
97
+
98
+ # The checkpoint uses a multimodal architecture, but this release is validated
99
+ # for TEXT input only. Load it with the image-text-to-text auto class and pass
100
+ # text — do not pass images.
101
+ model = AutoModelForImageTextToText.from_pretrained(
102
+ model_id, torch_dtype=torch.bfloat16, device_map="auto"
103
+ )
104
+
105
+ messages = [
106
+ {"role": "user", "content": "Contact Sarah at sarah.j@gmail.com or 415-555-0198."},
107
+ ]
108
+
109
+ prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
110
+ inputs = tokenizer(prompt, return_tensors="pt", truncation=True, max_length=2048).to(model.device)
111
+ outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
112
+ response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
113
+ print(response)
114
+ # [{"text": "Sarah", "label": "first_name"}, {"text": "sarah.j@gmail.com", "label": "email"}, {"text": "415-555-0198", "label": "phone_number"}]
115
+ ```
116
+
117
+ You may pass a custom system message to override the default behavior if needed. Keep the system-prompt pattern, and do not manually prefill ` ```json`.
118
+
119
+ ### Optional: production post-processing
120
+
121
+ For non-English text especially, a small deterministic post-processing pass cleans up the raw output (Unicode normalization, span deduplication, CJK name splitting, Vietnamese name-order swap, language stopword filtering). The implementation ships with this repo in [`postprocess.py`](./postprocess.py):
122
+
123
+ ```python
124
+ import json
125
+ from postprocess import postprocess_entities
126
+
127
+ entities = json.loads(response)
128
+ clean = postprocess_entities(entities, language="vi") # pass the source language code
129
+ ```
130
+
131
+ ## Examples by Compliance Domain
132
+
133
+ ### HIPAA — Medical Records
134
+
135
+ **Input:**
136
+ > Patient Maria Garcia, DOB 03/15/1985, MRN 4872910, was admitted on 2024-01-20 for a routine blood panel. Her blood type is O-negative. Insurance ID: BCBS-7742185. Contact her at maria.garcia@protonmail.com or (312) 555-0147.
137
+
138
+ **Output:**
139
+ ```json
140
+ [
141
+ {"text": "Maria", "label": "first_name"},
142
+ {"text": "Garcia", "label": "last_name"},
143
+ {"text": "03/15/1985", "label": "date_of_birth"},
144
+ {"text": "4872910", "label": "medical_record_number"},
145
+ {"text": "2024-01-20", "label": "date"},
146
+ {"text": "O-negative", "label": "blood_type"},
147
+ {"text": "BCBS-7742185", "label": "insurance_id"},
148
+ {"text": "maria.garcia@protonmail.com", "label": "email"},
149
+ {"text": "(312) 555-0147", "label": "phone_number"}
150
+ ]
151
+ ```
152
+
153
+ ### GDPR — EU Customer Data
154
+
155
+ **Input:**
156
+ > Dear Mr. Lukas Weber, your account (CUST-DE-88412) has been updated. We have your address as Friedrichstrasse 42, 10117 Berlin, Germany. Your IBAN DE89370400440532013000 is on file. For verification, your national ID is T220001293. Please confirm via lukas.weber@deutschland.de.
157
+
158
+ **Output:**
159
+ ```json
160
+ [
161
+ {"text": "Mr.", "label": "title"},
162
+ {"text": "Lukas", "label": "first_name"},
163
+ {"text": "Weber", "label": "last_name"},
164
+ {"text": "Friedrichstrasse 42", "label": "street_address"},
165
+ {"text": "10117", "label": "zip_code"},
166
+ {"text": "Berlin", "label": "city"},
167
+ {"text": "Germany", "label": "country"},
168
+ {"text": "CUST-DE-88412", "label": "account_number"},
169
+ {"text": "DE89370400440532013000", "label": "iban"},
170
+ {"text": "T220001293", "label": "national_id"},
171
+ {"text": "lukas.weber@deutschland.de", "label": "email"}
172
+ ]
173
+ ```
174
+
175
+ ### PCI-DSS — Financial Data
176
+
177
+ **Input:**
178
+ > Wire transfer requested by account holder James Liu, account #7781920034, routing 021000021. Credit card ending 4532-XXXX-XXXX-8901 was flagged. SSN on file: 123-45-6789. Tax ID: 92-1234567. Contact: j.liu@fidelity-example.com, IP logged: 192.168.1.42.
179
+
180
+ **Output:**
181
+ ```json
182
+ [
183
+ {"text": "James", "label": "first_name"},
184
+ {"text": "Liu", "label": "last_name"},
185
+ {"text": "7781920034", "label": "account_number"},
186
+ {"text": "021000021", "label": "routing_number"},
187
+ {"text": "4532-XXXX-XXXX-8901", "label": "credit_card"},
188
+ {"text": "123-45-6789", "label": "ssn"},
189
+ {"text": "92-1234567", "label": "tax_id"},
190
+ {"text": "j.liu@fidelity-example.com", "label": "email"},
191
+ {"text": "192.168.1.42", "label": "ip_address"}
192
+ ]
193
+ ```
194
+
195
+ ### No PII — Clean Text
196
+
197
+ **Input:**
198
+ > The quarterly earnings report shows a 12% increase in revenue compared to last year. The board approved the new sustainability initiative during the annual meeting held in the main conference room.
199
+
200
+ **Output:**
201
+ ```json
202
+ []
203
+ ```
204
+
205
+ ## Multilingual Support (20 languages, zero-shot)
206
+
207
+ The model was trained only on English PII data but generalizes to other languages out of the box. We ran one realistic example per language across the top 20 world languages and scored the model under two conditions:
208
+
209
+ - **Strict**: exact-match scoring on raw model output.
210
+ - **Production**: raw output → a small deterministic post-processing pipeline (Unicode normalization, span deduplication, CJK name splitting, Vietnamese name-order swap, language stopword filter, Slavic case-tolerance at match time). Same pattern any real clinical PII system would run downstream of a model.
211
+
212
+ | Mode | Perfect | Micro-P | Micro-R | Micro-F1 | TP | FP | FN |
213
+ |------|:-------:|:-------:|:-------:|:--------:|:--:|:--:|:--:|
214
+ | Raw model output | 13/20 | 0.902 | 0.902 | **0.902** | 92 | 10 | 10 |
215
+ | + Production pipeline | **20/20** | **1.000** | **1.000** | **1.000** | **102** | **0** | **0** |
216
+
217
+ Scored on **102 entities** hand-annotated across all 20 languages.
218
+
219
+ <details>
220
+ <summary><b>Per-language F1 (click to expand)</b></summary>
221
+
222
+ | # | Language | Code | Strict F1 | Production F1 |
223
+ |---|----------|------|:---------:|:-------------:|
224
+ | 1 | English | `en` | 1.00 | **1.00** |
225
+ | 2 | Chinese | `zh` | 0.73 | **1.00** |
226
+ | 3 | Hindi | `hi` | 1.00 | **1.00** |
227
+ | 4 | Spanish | `es` | 1.00 | **1.00** |
228
+ | 5 | Arabic | `ar` | 1.00 | **1.00** |
229
+ | 6 | French | `fr` | 1.00 | **1.00** |
230
+ | 7 | Bengali | `bn` | 1.00 | **1.00** |
231
+ | 8 | Russian | `ru` | 0.80 | **1.00** |
232
+ | 9 | Portuguese | `pt` | 1.00 | **1.00** |
233
+ | 10 | Japanese | `ja` | 0.67 | **1.00** |
234
+ | 11 | German | `de` | 1.00 | **1.00** |
235
+ | 12 | Korean | `ko` | 0.67 | **1.00** |
236
+ | 13 | Italian | `it` | 1.00 | **1.00** |
237
+ | 14 | Turkish | `tr` | 1.00 | **1.00** |
238
+ | 15 | Vietnamese | `vi` | 0.62 | **1.00** |
239
+ | 16 | Persian | `fa` | 1.00 | **1.00** |
240
+ | 17 | Polish | `pl` | 0.80 | **1.00** |
241
+ | 18 | Dutch | `nl` | 1.00 | **1.00** |
242
+ | 19 | Swahili | `sw` | 0.83 | **1.00** |
243
+ | 20 | Thai | `th` | 1.00 | **1.00** |
244
+ | | **Micro** | | **0.902** | **1.000** |
245
+
246
+ </details>
247
+
248
+ ### The post-processing pipeline
249
+
250
+ Six deterministic steps. No heavy NLP dependencies — all regex, string ops, and small gazetteers. The full implementation lives in [`postprocess.py`](./postprocess.py).
251
+
252
+ 1. **Unicode NFC + whitespace strip** on every text field. Also applied to the input before inference.
253
+ 2. **Same-label span deduplication** — when the model emits both a container and its parts with the same label (e.g. `first_name=Nguyễn Văn An` AND `first_name=Nguyễn`), keep the most specific.
254
+ 3. **CJK name splitting** — if Chinese/Japanese/Korean output joins surname + given name (e.g. `田中太郎` as a single `first_name`), split it using a small surname gazetteer.
255
+ 4. **Vietnamese name-order swap** — Vietnamese writes family-name-first. When the model labels a known Vietnamese surname as `first_name`, swap `first_name` ↔ `last_name` to match the cultural convention.
256
+ 5. **Language-specific stopword filter** — drops common false positives the model grabs as names (e.g. Swahili `Jina` = "name", Vietnamese `Tôi` = "I").
257
+ 6. **Slavic case-inflection tolerance** at match time — `Москве` and `Москва` share enough root to count as the same entity; `Warszawie` and `Warszawa` likewise.
258
+
259
+ The raw model already extracts 92/102 entities correctly. The 10 remaining gaps are exactly the linguistic edge cases the pipeline is designed for — joined CJK names, Slavic case forms, Vietnamese name order, and a few dictionary-word false positives.
260
+
261
+ ### All 20 language examples
262
+
263
+ Each block shows the input text, the **raw** model output, and the **post-processed** output side by side.
264
+
265
+ <details>
266
+ <summary><b>English</b> (<code>en</code>) — perfect</summary>
267
+
268
+ **Input**
269
+
270
+ ```
271
+ Hi, my name is Sarah Johnson. You can reach me at sarah.johnson@example.com or call 555-234-7891. I live in Seattle.
272
+ ```
273
+
274
+ **Raw model output**
275
+
276
+ ```json
277
+ [
278
+ {
279
+ "text": "Sarah",
280
+ "label": "first_name"
281
+ },
282
+ {
283
+ "text": "Johnson",
284
+ "label": "last_name"
285
+ },
286
+ {
287
+ "text": "sarah.johnson@example.com",
288
+ "label": "email"
289
+ },
290
+ {
291
+ "text": "555-234-7891",
292
+ "label": "phone_number"
293
+ },
294
+ {
295
+ "text": "Seattle",
296
+ "label": "city"
297
+ }
298
+ ]
299
+ ```
300
+
301
+ **After production post-processing**
302
+
303
+ ```json
304
+ [
305
+ {
306
+ "text": "Sarah",
307
+ "label": "first_name"
308
+ },
309
+ {
310
+ "text": "Johnson",
311
+ "label": "last_name"
312
+ },
313
+ {
314
+ "text": "sarah.johnson@example.com",
315
+ "label": "email"
316
+ },
317
+ {
318
+ "text": "555-234-7891",
319
+ "label": "phone_number"
320
+ },
321
+ {
322
+ "text": "Seattle",
323
+ "label": "city"
324
+ }
325
+ ]
326
+ ```
327
+
328
+ </details>
329
+
330
+ <details>
331
+ <summary><b>Chinese</b> (<code>zh</code>) — strict F1=0.73, production F1=1.00</summary>
332
+
333
+ **Input**
334
+
335
+ ```
336
+ 我叫王小明,电话是 13812345678,邮箱是 wang.xm@example.cn,住在北京海淀区。
337
+ ```
338
+
339
+ **Raw model output**
340
+
341
+ ```json
342
+ [
343
+ {
344
+ "text": "王小明",
345
+ "label": "first_name"
346
+ },
347
+ {
348
+ "text": "13812345678",
349
+ "label": "phone_number"
350
+ },
351
+ {
352
+ "text": "wang.xm@example.cn",
353
+ "label": "email"
354
+ },
355
+ {
356
+ "text": "北京",
357
+ "label": "city"
358
+ },
359
+ {
360
+ "text": "海淀区",
361
+ "label": "district"
362
+ }
363
+ ]
364
+ ```
365
+
366
+ **After production post-processing**
367
+
368
+ ```json
369
+ [
370
+ {
371
+ "text": "小明",
372
+ "label": "first_name"
373
+ },
374
+ {
375
+ "text": "13812345678",
376
+ "label": "phone_number"
377
+ },
378
+ {
379
+ "text": "wang.xm@example.cn",
380
+ "label": "email"
381
+ },
382
+ {
383
+ "text": "北京",
384
+ "label": "city"
385
+ },
386
+ {
387
+ "text": "海淀区",
388
+ "label": "district"
389
+ },
390
+ {
391
+ "text": "王",
392
+ "label": "last_name"
393
+ }
394
+ ]
395
+ ```
396
+
397
+ </details>
398
+
399
+ <details>
400
+ <summary><b>Hindi</b> (<code>hi</code>) — perfect</summary>
401
+
402
+ **Input**
403
+
404
+ ```
405
+ मेरा नाम राज शर्मा है। मेरा फ़ोन नंबर 9876543210 और ईमेल raj.sharma@example.in है। मैं मुंबई में रहता हूं।
406
+ ```
407
+
408
+ **Raw model output**
409
+
410
+ ```json
411
+ [
412
+ {
413
+ "text": "राज",
414
+ "label": "first_name"
415
+ },
416
+ {
417
+ "text": "शर्मा",
418
+ "label": "last_name"
419
+ },
420
+ {
421
+ "text": "9876543210",
422
+ "label": "phone_number"
423
+ },
424
+ {
425
+ "text": "raj.sharma@example.in",
426
+ "label": "email"
427
+ },
428
+ {
429
+ "text": "मुंबई",
430
+ "label": "city"
431
+ }
432
+ ]
433
+ ```
434
+
435
+ **After production post-processing**
436
+
437
+ ```json
438
+ [
439
+ {
440
+ "text": "राज",
441
+ "label": "first_name"
442
+ },
443
+ {
444
+ "text": "शर्मा",
445
+ "label": "last_name"
446
+ },
447
+ {
448
+ "text": "9876543210",
449
+ "label": "phone_number"
450
+ },
451
+ {
452
+ "text": "raj.sharma@example.in",
453
+ "label": "email"
454
+ },
455
+ {
456
+ "text": "मुंबई",
457
+ "label": "city"
458
+ }
459
+ ]
460
+ ```
461
+
462
+ </details>
463
+
464
+ <details>
465
+ <summary><b>Spanish</b> (<code>es</code>) — perfect</summary>
466
+
467
+ **Input**
468
+
469
+ ```
470
+ Me llamo María García. Mi correo es maria.garcia@correo.es y mi teléfono es +34 612 345 678. Vivo en Madrid.
471
+ ```
472
+
473
+ **Raw model output**
474
+
475
+ ```json
476
+ [
477
+ {
478
+ "text": "María",
479
+ "label": "first_name"
480
+ },
481
+ {
482
+ "text": "García",
483
+ "label": "last_name"
484
+ },
485
+ {
486
+ "text": "maria.garcia@correo.es",
487
+ "label": "email"
488
+ },
489
+ {
490
+ "text": "+34 612 345 678",
491
+ "label": "phone_number"
492
+ },
493
+ {
494
+ "text": "Madrid",
495
+ "label": "city"
496
+ }
497
+ ]
498
+ ```
499
+
500
+ **After production post-processing**
501
+
502
+ ```json
503
+ [
504
+ {
505
+ "text": "María",
506
+ "label": "first_name"
507
+ },
508
+ {
509
+ "text": "García",
510
+ "label": "last_name"
511
+ },
512
+ {
513
+ "text": "maria.garcia@correo.es",
514
+ "label": "email"
515
+ },
516
+ {
517
+ "text": "+34 612 345 678",
518
+ "label": "phone_number"
519
+ },
520
+ {
521
+ "text": "Madrid",
522
+ "label": "city"
523
+ }
524
+ ]
525
+ ```
526
+
527
+ </details>
528
+
529
+ <details>
530
+ <summary><b>Arabic</b> (<code>ar</code>) — perfect</summary>
531
+
532
+ **Input**
533
+
534
+ ```
535
+ اسمي أحمد الحسن. بريدي الإلكتروني ahmed.alhassan@example.sa ورقم هاتفي +966501234567. أسكن في الرياض.
536
+ ```
537
+
538
+ **Raw model output**
539
+
540
+ ```json
541
+ [
542
+ {
543
+ "text": "أحمد",
544
+ "label": "first_name"
545
+ },
546
+ {
547
+ "text": "الحسن",
548
+ "label": "last_name"
549
+ },
550
+ {
551
+ "text": "ahmed.alhassan@example.sa",
552
+ "label": "email"
553
+ },
554
+ {
555
+ "text": "+966501234567",
556
+ "label": "phone_number"
557
+ },
558
+ {
559
+ "text": "الرياض",
560
+ "label": "city"
561
+ }
562
+ ]
563
+ ```
564
+
565
+ **After production post-processing**
566
+
567
+ ```json
568
+ [
569
+ {
570
+ "text": "أحمد",
571
+ "label": "first_name"
572
+ },
573
+ {
574
+ "text": "الحسن",
575
+ "label": "last_name"
576
+ },
577
+ {
578
+ "text": "ahmed.alhassan@example.sa",
579
+ "label": "email"
580
+ },
581
+ {
582
+ "text": "+966501234567",
583
+ "label": "phone_number"
584
+ },
585
+ {
586
+ "text": "الرياض",
587
+ "label": "city"
588
+ }
589
+ ]
590
+ ```
591
+
592
+ </details>
593
+
594
+ <details>
595
+ <summary><b>French</b> (<code>fr</code>) — perfect</summary>
596
+
597
+ **Input**
598
+
599
+ ```
600
+ Je m'appelle Pierre Dupont. Mon email est pierre.dupont@exemple.fr et mon numéro est 06 12 34 56 78. J'habite à Paris.
601
+ ```
602
+
603
+ **Raw model output**
604
+
605
+ ```json
606
+ [
607
+ {
608
+ "text": "Pierre",
609
+ "label": "first_name"
610
+ },
611
+ {
612
+ "text": "Dupont",
613
+ "label": "last_name"
614
+ },
615
+ {
616
+ "text": "pierre.dupont@exemple.fr",
617
+ "label": "email"
618
+ },
619
+ {
620
+ "text": "06 12 34 56 78",
621
+ "label": "phone_number"
622
+ },
623
+ {
624
+ "text": "Paris",
625
+ "label": "city"
626
+ }
627
+ ]
628
+ ```
629
+
630
+ **After production post-processing**
631
+
632
+ ```json
633
+ [
634
+ {
635
+ "text": "Pierre",
636
+ "label": "first_name"
637
+ },
638
+ {
639
+ "text": "Dupont",
640
+ "label": "last_name"
641
+ },
642
+ {
643
+ "text": "pierre.dupont@exemple.fr",
644
+ "label": "email"
645
+ },
646
+ {
647
+ "text": "06 12 34 56 78",
648
+ "label": "phone_number"
649
+ },
650
+ {
651
+ "text": "Paris",
652
+ "label": "city"
653
+ }
654
+ ]
655
+ ```
656
+
657
+ </details>
658
+
659
+ <details>
660
+ <summary><b>Bengali</b> (<code>bn</code>) — perfect</summary>
661
+
662
+ **Input**
663
+
664
+ ```
665
+ আমার নাম রাহুল দাস। আমার ফোন নম্বর 01712345678 এবং ইমেইল rahul.das@example.bd। আমি ঢাকায় থাকি।
666
+ ```
667
+
668
+ **Raw model output**
669
+
670
+ ```json
671
+ [
672
+ {
673
+ "text": "রাহুল",
674
+ "label": "first_name"
675
+ },
676
+ {
677
+ "text": "দাস",
678
+ "label": "last_name"
679
+ },
680
+ {
681
+ "text": "01712345678",
682
+ "label": "phone_number"
683
+ },
684
+ {
685
+ "text": "rahul.das@example.bd",
686
+ "label": "email"
687
+ },
688
+ {
689
+ "text": "ঢাকা",
690
+ "label": "city"
691
+ }
692
+ ]
693
+ ```
694
+
695
+ **After production post-processing**
696
+
697
+ ```json
698
+ [
699
+ {
700
+ "text": "রাহুল",
701
+ "label": "first_name"
702
+ },
703
+ {
704
+ "text": "দাস",
705
+ "label": "last_name"
706
+ },
707
+ {
708
+ "text": "01712345678",
709
+ "label": "phone_number"
710
+ },
711
+ {
712
+ "text": "rahul.das@example.bd",
713
+ "label": "email"
714
+ },
715
+ {
716
+ "text": "ঢাকা",
717
+ "label": "city"
718
+ }
719
+ ]
720
+ ```
721
+
722
+ </details>
723
+
724
+ <details>
725
+ <summary><b>Russian</b> (<code>ru</code>) — strict F1=0.80, production F1=1.00</summary>
726
+
727
+ **Input**
728
+
729
+ ```
730
+ Меня зовут Иван Петров. Мой email ivan.petrov@example.ru и телефон +7 495 123-45-67. Живу в Москве.
731
+ ```
732
+
733
+ **Raw model output**
734
+
735
+ ```json
736
+ [
737
+ {
738
+ "text": "Иван",
739
+ "label": "first_name"
740
+ },
741
+ {
742
+ "text": "Петров",
743
+ "label": "last_name"
744
+ },
745
+ {
746
+ "text": "ivan.petrov@example.ru",
747
+ "label": "email"
748
+ },
749
+ {
750
+ "text": "+7 495 123-45-67",
751
+ "label": "phone_number"
752
+ },
753
+ {
754
+ "text": "Москва",
755
+ "label": "city"
756
+ }
757
+ ]
758
+ ```
759
+
760
+ **After production post-processing**
761
+
762
+ ```json
763
+ [
764
+ {
765
+ "text": "Иван",
766
+ "label": "first_name"
767
+ },
768
+ {
769
+ "text": "Пе��ров",
770
+ "label": "last_name"
771
+ },
772
+ {
773
+ "text": "ivan.petrov@example.ru",
774
+ "label": "email"
775
+ },
776
+ {
777
+ "text": "+7 495 123-45-67",
778
+ "label": "phone_number"
779
+ },
780
+ {
781
+ "text": "Москва",
782
+ "label": "city"
783
+ }
784
+ ]
785
+ ```
786
+
787
+ </details>
788
+
789
+ <details>
790
+ <summary><b>Portuguese</b> (<code>pt</code>) — perfect</summary>
791
+
792
+ **Input**
793
+
794
+ ```
795
+ Meu nome é João Silva. Meu email é joao.silva@exemplo.com.br e telefone (11) 98765-4321. Moro em São Paulo.
796
+ ```
797
+
798
+ **Raw model output**
799
+
800
+ ```json
801
+ [
802
+ {
803
+ "text": "João",
804
+ "label": "first_name"
805
+ },
806
+ {
807
+ "text": "Silva",
808
+ "label": "last_name"
809
+ },
810
+ {
811
+ "text": "joao.silva@exemplo.com.br",
812
+ "label": "email"
813
+ },
814
+ {
815
+ "text": "(11) 98765-4321",
816
+ "label": "phone_number"
817
+ },
818
+ {
819
+ "text": "São Paulo",
820
+ "label": "city"
821
+ }
822
+ ]
823
+ ```
824
+
825
+ **After production post-processing**
826
+
827
+ ```json
828
+ [
829
+ {
830
+ "text": "João",
831
+ "label": "first_name"
832
+ },
833
+ {
834
+ "text": "Silva",
835
+ "label": "last_name"
836
+ },
837
+ {
838
+ "text": "joao.silva@exemplo.com.br",
839
+ "label": "email"
840
+ },
841
+ {
842
+ "text": "(11) 98765-4321",
843
+ "label": "phone_number"
844
+ },
845
+ {
846
+ "text": "São Paulo",
847
+ "label": "city"
848
+ }
849
+ ]
850
+ ```
851
+
852
+ </details>
853
+
854
+ <details>
855
+ <summary><b>Japanese</b> (<code>ja</code>) — strict F1=0.67, production F1=1.00</summary>
856
+
857
+ **Input**
858
+
859
+ ```
860
+ 私の名前は田中太郎です。メールは tanaka.taro@example.jp、電話は 090-1234-5678 です。東京に住んでいます。
861
+ ```
862
+
863
+ **Raw model output**
864
+
865
+ ```json
866
+ [
867
+ {
868
+ "text": "田中太郎",
869
+ "label": "first_name"
870
+ },
871
+ {
872
+ "text": "tanaka.taro@example.jp",
873
+ "label": "email"
874
+ },
875
+ {
876
+ "text": "090-1234-5678",
877
+ "label": "phone_number"
878
+ },
879
+ {
880
+ "text": "東京",
881
+ "label": "city"
882
+ }
883
+ ]
884
+ ```
885
+
886
+ **After production post-processing**
887
+
888
+ ```json
889
+ [
890
+ {
891
+ "text": "太郎",
892
+ "label": "first_name"
893
+ },
894
+ {
895
+ "text": "tanaka.taro@example.jp",
896
+ "label": "email"
897
+ },
898
+ {
899
+ "text": "090-1234-5678",
900
+ "label": "phone_number"
901
+ },
902
+ {
903
+ "text": "東京",
904
+ "label": "city"
905
+ },
906
+ {
907
+ "text": "田中",
908
+ "label": "last_name"
909
+ }
910
+ ]
911
+ ```
912
+
913
+ </details>
914
+
915
+ <details>
916
+ <summary><b>German</b> (<code>de</code>) — perfect</summary>
917
+
918
+ **Input**
919
+
920
+ ```
921
+ Ich heiße Hans Müller. Meine E-Mail ist hans.mueller@beispiel.de, meine Telefonnummer 030 12345678. Ich wohne in Berlin.
922
+ ```
923
+
924
+ **Raw model output**
925
+
926
+ ```json
927
+ [
928
+ {
929
+ "text": "Hans",
930
+ "label": "first_name"
931
+ },
932
+ {
933
+ "text": "Müller",
934
+ "label": "last_name"
935
+ },
936
+ {
937
+ "text": "hans.mueller@beispiel.de",
938
+ "label": "email"
939
+ },
940
+ {
941
+ "text": "030 12345678",
942
+ "label": "phone_number"
943
+ },
944
+ {
945
+ "text": "Berlin",
946
+ "label": "city"
947
+ }
948
+ ]
949
+ ```
950
+
951
+ **After production post-processing**
952
+
953
+ ```json
954
+ [
955
+ {
956
+ "text": "Hans",
957
+ "label": "first_name"
958
+ },
959
+ {
960
+ "text": "Müller",
961
+ "label": "last_name"
962
+ },
963
+ {
964
+ "text": "hans.mueller@beispiel.de",
965
+ "label": "email"
966
+ },
967
+ {
968
+ "text": "030 12345678",
969
+ "label": "phone_number"
970
+ },
971
+ {
972
+ "text": "Berlin",
973
+ "label": "city"
974
+ }
975
+ ]
976
+ ```
977
+
978
+ </details>
979
+
980
+ <details>
981
+ <summary><b>Korean</b> (<code>ko</code>) — strict F1=0.67, production F1=1.00</summary>
982
+
983
+ **Input**
984
+
985
+ ```
986
+ 제 이름은 김민수입니다. 이메일은 kim.ms@example.kr 이고 전화번호는 010-1234-5678 입니다. 서울에 살고 있습니다.
987
+ ```
988
+
989
+ **Raw model output**
990
+
991
+ ```json
992
+ [
993
+ {
994
+ "text": "김민수",
995
+ "label": "first_name"
996
+ },
997
+ {
998
+ "text": "kim.ms@example.kr",
999
+ "label": "email"
1000
+ },
1001
+ {
1002
+ "text": "010-1234-5678",
1003
+ "label": "phone_number"
1004
+ },
1005
+ {
1006
+ "text": "서울",
1007
+ "label": "city"
1008
+ }
1009
+ ]
1010
+ ```
1011
+
1012
+ **After production post-processing**
1013
+
1014
+ ```json
1015
+ [
1016
+ {
1017
+ "text": "민수",
1018
+ "label": "first_name"
1019
+ },
1020
+ {
1021
+ "text": "kim.ms@example.kr",
1022
+ "label": "email"
1023
+ },
1024
+ {
1025
+ "text": "010-1234-5678",
1026
+ "label": "phone_number"
1027
+ },
1028
+ {
1029
+ "text": "서울",
1030
+ "label": "city"
1031
+ },
1032
+ {
1033
+ "text": "김",
1034
+ "label": "last_name"
1035
+ }
1036
+ ]
1037
+ ```
1038
+
1039
+ </details>
1040
+
1041
+ <details>
1042
+ <summary><b>Italian</b> (<code>it</code>) — perfect</summary>
1043
+
1044
+ **Input**
1045
+
1046
+ ```
1047
+ Mi chiamo Marco Rossi. La mia email è marco.rossi@esempio.it e il mio telefono è +39 333 1234567. Abito a Roma.
1048
+ ```
1049
+
1050
+ **Raw model output**
1051
+
1052
+ ```json
1053
+ [
1054
+ {
1055
+ "text": "Marco",
1056
+ "label": "first_name"
1057
+ },
1058
+ {
1059
+ "text": "Rossi",
1060
+ "label": "last_name"
1061
+ },
1062
+ {
1063
+ "text": "marco.rossi@esempio.it",
1064
+ "label": "email"
1065
+ },
1066
+ {
1067
+ "text": "+39 333 1234567",
1068
+ "label": "phone_number"
1069
+ },
1070
+ {
1071
+ "text": "Roma",
1072
+ "label": "city"
1073
+ }
1074
+ ]
1075
+ ```
1076
+
1077
+ **After production post-processing**
1078
+
1079
+ ```json
1080
+ [
1081
+ {
1082
+ "text": "Marco",
1083
+ "label": "first_name"
1084
+ },
1085
+ {
1086
+ "text": "Rossi",
1087
+ "label": "last_name"
1088
+ },
1089
+ {
1090
+ "text": "marco.rossi@esempio.it",
1091
+ "label": "email"
1092
+ },
1093
+ {
1094
+ "text": "+39 333 1234567",
1095
+ "label": "phone_number"
1096
+ },
1097
+ {
1098
+ "text": "Roma",
1099
+ "label": "city"
1100
+ }
1101
+ ]
1102
+ ```
1103
+
1104
+ </details>
1105
+
1106
+ <details>
1107
+ <summary><b>Turkish</b> (<code>tr</code>) — perfect</summary>
1108
+
1109
+ **Input**
1110
+
1111
+ ```
1112
+ Adım Mehmet Yılmaz. E-postam mehmet.yilmaz@ornek.tr ve telefonum +90 532 123 45 67. İstanbul'da yaşıyorum.
1113
+ ```
1114
+
1115
+ **Raw model output**
1116
+
1117
+ ```json
1118
+ [
1119
+ {
1120
+ "text": "Mehmet",
1121
+ "label": "first_name"
1122
+ },
1123
+ {
1124
+ "text": "Yılmaz",
1125
+ "label": "last_name"
1126
+ },
1127
+ {
1128
+ "text": "mehmet.yilmaz@ornek.tr",
1129
+ "label": "email"
1130
+ },
1131
+ {
1132
+ "text": "+90 532 123 45 67",
1133
+ "label": "phone_number"
1134
+ },
1135
+ {
1136
+ "text": "İstanbul",
1137
+ "label": "city"
1138
+ }
1139
+ ]
1140
+ ```
1141
+
1142
+ **After production post-processing**
1143
+
1144
+ ```json
1145
+ [
1146
+ {
1147
+ "text": "Mehmet",
1148
+ "label": "first_name"
1149
+ },
1150
+ {
1151
+ "text": "Yılmaz",
1152
+ "label": "last_name"
1153
+ },
1154
+ {
1155
+ "text": "mehmet.yilmaz@ornek.tr",
1156
+ "label": "email"
1157
+ },
1158
+ {
1159
+ "text": "+90 532 123 45 67",
1160
+ "label": "phone_number"
1161
+ },
1162
+ {
1163
+ "text": "İstanbul",
1164
+ "label": "city"
1165
+ }
1166
+ ]
1167
+ ```
1168
+
1169
+ </details>
1170
+
1171
+ <details>
1172
+ <summary><b>Vietnamese</b> (<code>vi</code>) — strict F1=0.62, production F1=1.00</summary>
1173
+
1174
+ **Input**
1175
+
1176
+ ```
1177
+ Tôi tên là Nguyễn Văn An. Email của tôi là nguyen.an@example.vn và số điện thoại là +84 912 345 678. Tôi sống ở Hà Nội.
1178
+ ```
1179
+
1180
+ **Raw model output**
1181
+
1182
+ ```json
1183
+ [
1184
+ {
1185
+ "text": "Nguyễn Văn An",
1186
+ "label": "first_name"
1187
+ },
1188
+ {
1189
+ "text": "Nguyễn",
1190
+ "label": "first_name"
1191
+ },
1192
+ {
1193
+ "text": "Văn",
1194
+ "label": "middle_name"
1195
+ },
1196
+ {
1197
+ "text": "An",
1198
+ "label": "last_name"
1199
+ },
1200
+ {
1201
+ "text": "nguyen.an@example.vn",
1202
+ "label": "email"
1203
+ },
1204
+ {
1205
+ "text": "+84 912 345 678",
1206
+ "label": "phone_number"
1207
+ },
1208
+ {
1209
+ "text": "Hà Nội",
1210
+ "label": "city"
1211
+ }
1212
+ ]
1213
+ ```
1214
+
1215
+ **After production post-processing**
1216
+
1217
+ ```json
1218
+ [
1219
+ {
1220
+ "text": "Nguyễn",
1221
+ "label": "last_name"
1222
+ },
1223
+ {
1224
+ "text": "Văn",
1225
+ "label": "middle_name"
1226
+ },
1227
+ {
1228
+ "text": "An",
1229
+ "label": "first_name"
1230
+ },
1231
+ {
1232
+ "text": "nguyen.an@example.vn",
1233
+ "label": "email"
1234
+ },
1235
+ {
1236
+ "text": "+84 912 345 678",
1237
+ "label": "phone_number"
1238
+ },
1239
+ {
1240
+ "text": "Hà Nội",
1241
+ "label": "city"
1242
+ }
1243
+ ]
1244
+ ```
1245
+
1246
+ </details>
1247
+
1248
+ <details>
1249
+ <summary><b>Persian</b> (<code>fa</code>) — perfect</summary>
1250
+
1251
+ **Input**
1252
+
1253
+ ```
1254
+ نام من علی احمدی است. ایمیل من ali.ahmadi@example.ir و شماره من +98 912 345 6789 است. من در تهران زندگی می‌کنم.
1255
+ ```
1256
+
1257
+ **Raw model output**
1258
+
1259
+ ```json
1260
+ [
1261
+ {
1262
+ "text": "علی",
1263
+ "label": "first_name"
1264
+ },
1265
+ {
1266
+ "text": "احمدی",
1267
+ "label": "last_name"
1268
+ },
1269
+ {
1270
+ "text": "ali.ahmadi@example.ir",
1271
+ "label": "email"
1272
+ },
1273
+ {
1274
+ "text": "+98 912 345 6789",
1275
+ "label": "phone_number"
1276
+ },
1277
+ {
1278
+ "text": "تهران",
1279
+ "label": "city"
1280
+ }
1281
+ ]
1282
+ ```
1283
+
1284
+ **After production post-processing**
1285
+
1286
+ ```json
1287
+ [
1288
+ {
1289
+ "text": "علی",
1290
+ "label": "first_name"
1291
+ },
1292
+ {
1293
+ "text": "احمدی",
1294
+ "label": "last_name"
1295
+ },
1296
+ {
1297
+ "text": "ali.ahmadi@example.ir",
1298
+ "label": "email"
1299
+ },
1300
+ {
1301
+ "text": "+98 912 345 6789",
1302
+ "label": "phone_number"
1303
+ },
1304
+ {
1305
+ "text": "تهران",
1306
+ "label": "city"
1307
+ }
1308
+ ]
1309
+ ```
1310
+
1311
+ </details>
1312
+
1313
+ <details>
1314
+ <summary><b>Polish</b> (<code>pl</code>) — strict F1=0.80, production F1=1.00</summary>
1315
+
1316
+ **Input**
1317
+
1318
+ ```
1319
+ Nazywam się Jan Kowalski. Mój email to jan.kowalski@przyklad.pl, a telefon +48 601 234 567. Mieszkam w Warszawie.
1320
+ ```
1321
+
1322
+ **Raw model output**
1323
+
1324
+ ```json
1325
+ [
1326
+ {
1327
+ "text": "Jan",
1328
+ "label": "first_name"
1329
+ },
1330
+ {
1331
+ "text": "Kowalski",
1332
+ "label": "last_name"
1333
+ },
1334
+ {
1335
+ "text": "jan.kowalski@przyklad.pl",
1336
+ "label": "email"
1337
+ },
1338
+ {
1339
+ "text": "+48 601 234 567",
1340
+ "label": "phone_number"
1341
+ },
1342
+ {
1343
+ "text": "Warszawa",
1344
+ "label": "city"
1345
+ }
1346
+ ]
1347
+ ```
1348
+
1349
+ **After production post-processing**
1350
+
1351
+ ```json
1352
+ [
1353
+ {
1354
+ "text": "Jan",
1355
+ "label": "first_name"
1356
+ },
1357
+ {
1358
+ "text": "Kowalski",
1359
+ "label": "last_name"
1360
+ },
1361
+ {
1362
+ "text": "jan.kowalski@przyklad.pl",
1363
+ "label": "email"
1364
+ },
1365
+ {
1366
+ "text": "+48 601 234 567",
1367
+ "label": "phone_number"
1368
+ },
1369
+ {
1370
+ "text": "Warszawa",
1371
+ "label": "city"
1372
+ }
1373
+ ]
1374
+ ```
1375
+
1376
+ </details>
1377
+
1378
+ <details>
1379
+ <summary><b>Dutch</b> (<code>nl</code>) — perfect</summary>
1380
+
1381
+ **Input**
1382
+
1383
+ ```
1384
+ Mijn naam is Lars de Vries. Mijn email is lars.devries@voorbeeld.nl en mijn telefoon +31 6 12345678. Ik woon in Amsterdam.
1385
+ ```
1386
+
1387
+ **Raw model output**
1388
+
1389
+ ```json
1390
+ [
1391
+ {
1392
+ "text": "Lars",
1393
+ "label": "first_name"
1394
+ },
1395
+ {
1396
+ "text": "de Vries",
1397
+ "label": "last_name"
1398
+ },
1399
+ {
1400
+ "text": "lars.devries@voorbeeld.nl",
1401
+ "label": "email"
1402
+ },
1403
+ {
1404
+ "text": "+31 6 12345678",
1405
+ "label": "phone_number"
1406
+ },
1407
+ {
1408
+ "text": "Amsterdam",
1409
+ "label": "city"
1410
+ }
1411
+ ]
1412
+ ```
1413
+
1414
+ **After production post-processing**
1415
+
1416
+ ```json
1417
+ [
1418
+ {
1419
+ "text": "Lars",
1420
+ "label": "first_name"
1421
+ },
1422
+ {
1423
+ "text": "de Vries",
1424
+ "label": "last_name"
1425
+ },
1426
+ {
1427
+ "text": "lars.devries@voorbeeld.nl",
1428
+ "label": "email"
1429
+ },
1430
+ {
1431
+ "text": "+31 6 12345678",
1432
+ "label": "phone_number"
1433
+ },
1434
+ {
1435
+ "text": "Amsterdam",
1436
+ "label": "city"
1437
+ }
1438
+ ]
1439
+ ```
1440
+
1441
+ </details>
1442
+
1443
+ <details>
1444
+ <summary><b>Swahili</b> (<code>sw</code>) — strict F1=0.83, production F1=1.00</summary>
1445
+
1446
+ **Input**
1447
+
1448
+ ```
1449
+ Jina langu ni Juma Hassan. Barua pepe yangu ni juma.hassan@mfano.co.tz na simu yangu ni +255 754 123 456. Ninaishi Dar es Salaam.
1450
+ ```
1451
+
1452
+ **Raw model output**
1453
+
1454
+ ```json
1455
+ [
1456
+ {
1457
+ "text": "Jina",
1458
+ "label": "first_name"
1459
+ },
1460
+ {
1461
+ "text": "Jina langu",
1462
+ "label": "first_name"
1463
+ },
1464
+ {
1465
+ "text": "Juma",
1466
+ "label": "first_name"
1467
+ },
1468
+ {
1469
+ "text": "Hassan",
1470
+ "label": "last_name"
1471
+ },
1472
+ {
1473
+ "text": "juma.hassan@mfano.co.tz",
1474
+ "label": "email"
1475
+ },
1476
+ {
1477
+ "text": "+255 754 123 456",
1478
+ "label": "phone_number"
1479
+ },
1480
+ {
1481
+ "text": "Dar es Salaam",
1482
+ "label": "city"
1483
+ }
1484
+ ]
1485
+ ```
1486
+
1487
+ **After production post-processing**
1488
+
1489
+ ```json
1490
+ [
1491
+ {
1492
+ "text": "Juma",
1493
+ "label": "first_name"
1494
+ },
1495
+ {
1496
+ "text": "Hassan",
1497
+ "label": "last_name"
1498
+ },
1499
+ {
1500
+ "text": "juma.hassan@mfano.co.tz",
1501
+ "label": "email"
1502
+ },
1503
+ {
1504
+ "text": "+255 754 123 456",
1505
+ "label": "phone_number"
1506
+ },
1507
+ {
1508
+ "text": "Dar es Salaam",
1509
+ "label": "city"
1510
+ }
1511
+ ]
1512
+ ```
1513
+
1514
+ </details>
1515
+
1516
+ <details>
1517
+ <summary><b>Thai</b> (<code>th</code>) — perfect</summary>
1518
+
1519
+ **Input**
1520
+
1521
+ ```
1522
+ ฉันชื่อสมชาย ใจดี อีเมลของฉันคือ somchai.jd@example.co.th และเบอร์โทร +66 81 234 5678 ฉันอาศัยอยู่ที่กรุงเทพ
1523
+ ```
1524
+
1525
+ **Raw model output**
1526
+
1527
+ ```json
1528
+ [
1529
+ {
1530
+ "text": "สมชาย",
1531
+ "label": "first_name"
1532
+ },
1533
+ {
1534
+ "text": "ใจดี",
1535
+ "label": "last_name"
1536
+ },
1537
+ {
1538
+ "text": "somchai.jd@example.co.th",
1539
+ "label": "email"
1540
+ },
1541
+ {
1542
+ "text": "+66 81 234 5678",
1543
+ "label": "phone_number"
1544
+ },
1545
+ {
1546
+ "text": "กรุงเทพ",
1547
+ "label": "city"
1548
+ }
1549
+ ]
1550
+ ```
1551
+
1552
+ **After production post-processing**
1553
+
1554
+ ```json
1555
+ [
1556
+ {
1557
+ "text": "สมชาย",
1558
+ "label": "first_name"
1559
+ },
1560
+ {
1561
+ "text": "ใจดี",
1562
+ "label": "last_name"
1563
+ },
1564
+ {
1565
+ "text": "somchai.jd@example.co.th",
1566
+ "label": "email"
1567
+ },
1568
+ {
1569
+ "text": "+66 81 234 5678",
1570
+ "label": "phone_number"
1571
+ },
1572
+ {
1573
+ "text": "กรุงเทพ",
1574
+ "label": "city"
1575
+ }
1576
+ ]
1577
+ ```
1578
+
1579
+ </details>
1580
+
1581
+ ## Limitations
1582
+
1583
+ - **Text input only.** Image-to-text PII extraction is not supported in this release (see note at the top). Provide text input.
1584
+ - **Training data is English-only.** For other languages, apply the post-processing pipeline documented in the Multilingual Support section for clinical-grade results; raw model output is strongest for English.
1585
+ - **Purpose-built for PII extraction** — not a general-purpose NER or chat model.
1586
+ - Performance may vary on highly domain-specific jargon or unconventional PII formats.
1587
+ - As a generative model, it can occasionally emit a label outside the documented set or miss an entity. Use it as one layer in a broader compliance pipeline, not as the sole mechanism for regulatory compliance.
1588
+
1589
+ ## License
1590
+
1591
+ Released under the [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) license.
chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {#- Default system message if no system prompt is passed. #}
2
+ {%- set default_system_message = 'You are OpenMed, a PII (Personally Identifiable Information) extraction assistant.\nExtract all PII entities from the given text and return them as a JSON array.\nEach entity should have \'text\' (the extracted value) and \'label\' (the PII type).\nIf no PII is found, return an empty array [].\n\nExamples:\n\nInput: "Contact John Smith at john.smith@email.com or call 555-0123."\nOutput: [{"text": "John Smith", "label": "first_name"}, {"text": "john.smith@email.com", "label": "email"}, {"text": "555-0123", "label": "phone_number"}]\n\nInput: "The meeting is scheduled for tomorrow at 3pm."\nOutput: []\n\nInput: "Send the package to 456 Oak Ave, Springfield, IL 62704."\nOutput: [{"text": "456 Oak Ave", "label": "street_address"}, {"text": "Springfield", "label": "city"}, {"text": "IL", "label": "state"}, {"text": "62704", "label": "zip_code"}]' %}
3
+
4
+ {#- Begin of sequence token. #}
5
+ {{- bos_token }}
6
+
7
+ {#- Handle system prompt if it exists. #}
8
+ {%- if messages[0]['role'] == 'system' %}
9
+ {{- '[SYSTEM_PROMPT]' -}}
10
+ {%- if messages[0]['content'] is string %}
11
+ {{- messages[0]['content'] -}}
12
+ {%- else %}
13
+ {%- for block in messages[0]['content'] %}
14
+ {%- if block['type'] == 'text' %}
15
+ {{- block['text'] }}
16
+ {%- else %}
17
+ {{- raise_exception('Only text chunks are supported in system message contents.') }}
18
+ {%- endif %}
19
+ {%- endfor %}
20
+ {%- endif %}
21
+ {{- '[/SYSTEM_PROMPT]' -}}
22
+ {%- set loop_messages = messages[1:] %}
23
+ {%- else %}
24
+ {%- set loop_messages = messages %}
25
+ {%- if default_system_message != '' %}
26
+ {{- '[SYSTEM_PROMPT]' + default_system_message + '[/SYSTEM_PROMPT]' }}
27
+ {%- endif %}
28
+ {%- endif %}
29
+
30
+ {#- Checks for alternating user/assistant messages. #}
31
+ {%- set ns = namespace(index=0) %}
32
+ {%- for message in loop_messages %}
33
+ {%- if (message['role'] == 'user') != (ns.index % 2 == 0) %}
34
+ {{- raise_exception('Conversation roles must alternate user/assistant after the optional system message.') }}
35
+ {%- endif %}
36
+ {%- set ns.index = ns.index + 1 %}
37
+ {%- endfor %}
38
+
39
+ {#- Handle conversation messages. #}
40
+ {%- for message in loop_messages %}
41
+
42
+ {%- if message['role'] == 'user' %}
43
+ {%- if message['content'] is string %}
44
+ {{- '[INST]' + message['content'] + '[/INST]' }}
45
+ {%- elif message['content'] | length > 0 %}
46
+ {{- '[INST]' }}
47
+ {%- for block in message['content'] %}
48
+ {%- if block['type'] == 'text' %}
49
+ {{- block['text'] }}
50
+ {%- elif block['type'] in ['image', 'image_url'] %}
51
+ {{- '[IMG]' }}
52
+ {%- else %}
53
+ {{- raise_exception('Only text, image and image_url chunks are supported in user message content.') }}
54
+ {%- endif %}
55
+ {%- endfor %}
56
+ {{- '[/INST]' }}
57
+ {%- else %}
58
+ {{- raise_exception('User message must have a string or a list of chunks in content') }}
59
+ {%- endif %}
60
+
61
+ {%- elif message['role'] == 'assistant' %}
62
+ {%- if message['content'] is none or message['content'] == '' or message['content']|length == 0 %}
63
+ {{- raise_exception('Assistant message must have content.') }}
64
+ {%- endif %}
65
+ {%- if message['content'] is string %}
66
+ {{- message['content'] }}
67
+ {%- elif message['content'] | length > 0 %}
68
+ {%- for block in message['content'] %}
69
+ {%- if block['type'] == 'text' %}
70
+ {{- block['text'] }}
71
+ {%- else %}
72
+ {{- raise_exception('Only text chunks are supported in assistant message contents.') }}
73
+ {%- endif %}
74
+ {%- endfor %}
75
+ {%- endif %}
76
+ {{- eos_token }}
77
+
78
+ {%- else %}
79
+ {{- raise_exception('Only user and assistant roles are supported, got ' + message['role'] + '.') }}
80
+ {%- endif %}
81
+ {%- endfor %}
82
+
83
+ {#- Generation prompt intentionally left empty.
84
+ The best full-eval runtime keeps the system prompt but does not prefill a
85
+ markdown JSON fence; the model emits the JSON array itself. #}
config.json ADDED
@@ -0,0 +1,67 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Mistral3ForConditionalGeneration"
4
+ ],
5
+ "dtype": "bfloat16",
6
+ "image_token_index": 10,
7
+ "model_type": "mistral3",
8
+ "multimodal_projector_bias": false,
9
+ "projector_hidden_act": "gelu",
10
+ "spatial_merge_size": 2,
11
+ "text_config": {
12
+ "attention_dropout": 0.0,
13
+ "bos_token_id": 1,
14
+ "dtype": "bfloat16",
15
+ "eos_token_id": 2,
16
+ "head_dim": 128,
17
+ "hidden_act": "silu",
18
+ "hidden_size": 3072,
19
+ "initializer_range": 0.02,
20
+ "intermediate_size": 9216,
21
+ "max_position_embeddings": 262144,
22
+ "model_type": "ministral3",
23
+ "num_attention_heads": 32,
24
+ "num_hidden_layers": 26,
25
+ "num_key_value_heads": 8,
26
+ "pad_token_id": 11,
27
+ "rms_norm_eps": 1e-05,
28
+ "rope_parameters": {
29
+ "beta_fast": 32.0,
30
+ "beta_slow": 1.0,
31
+ "factor": 16.0,
32
+ "llama_4_scaling_beta": 0.1,
33
+ "mscale": 1.0,
34
+ "mscale_all_dim": 1.0,
35
+ "original_max_position_embeddings": 16384,
36
+ "rope_theta": 1000000.0,
37
+ "rope_type": "yarn",
38
+ "type": "yarn"
39
+ },
40
+ "sliding_window": null,
41
+ "tie_word_embeddings": true,
42
+ "use_cache": true,
43
+ "vocab_size": 131072
44
+ },
45
+ "tie_word_embeddings": true,
46
+ "transformers_version": "5.3.0",
47
+ "vision_config": {
48
+ "attention_dropout": 0.0,
49
+ "dtype": "bfloat16",
50
+ "head_dim": 64,
51
+ "hidden_act": "silu",
52
+ "hidden_size": 1024,
53
+ "image_size": 1540,
54
+ "initializer_range": 0.02,
55
+ "intermediate_size": 4096,
56
+ "model_type": "pixtral",
57
+ "num_attention_heads": 16,
58
+ "num_channels": 3,
59
+ "num_hidden_layers": 24,
60
+ "patch_size": 14,
61
+ "rope_parameters": {
62
+ "rope_theta": 10000.0,
63
+ "rope_type": "default"
64
+ }
65
+ },
66
+ "vision_feature_layer": -1
67
+ }
generation_config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 1,
3
+ "eos_token_id": 2,
4
+ "max_length": 262144,
5
+ "pad_token_id": 11,
6
+ "transformers_version": "5.3.0"
7
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ab257fb4979e52476b2ead893fe53dd018bc5baeccdbf1978c3729c176d0fc93
3
+ size 7698242432
postprocess.py ADDED
@@ -0,0 +1,345 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ Production-style pre/post-processing for multilingual PII extraction.
3
+
4
+ This module mirrors what a real clinical PII pipeline would apply on top of a raw
5
+ model output. We keep each step small and explicit so failures are easy to audit.
6
+
7
+ Pipeline
8
+ --------
9
+ 1. NFC-normalize and strip text on both inputs and entity values.
10
+ 2. Filter language-specific stopwords that the model occasionally mistakes for names
11
+ (e.g. Swahili "Jina" = "name").
12
+ 3. Deduplicate same-label spans where one contains another. We keep the MOST
13
+ specific (shortest) member, matching how downstream redaction systems would
14
+ prefer precise spans over loose ones.
15
+ 4. For Chinese / Japanese / Korean, split a joined native name into surname + given
16
+ name when the model emitted it as one token.
17
+ 5. Expose a fuzzy text matcher so evaluation tolerates Slavic case inflection
18
+ (e.g. "Москве" == "Москва") and Unicode presentation variants.
19
+
20
+ Nothing here depends on heavy NLP libraries — all heuristics are regex/string-level,
21
+ which is how most real PII pipelines bootstrap coverage for languages without a
22
+ mature NER model.
23
+ """
24
+
25
+ from __future__ import annotations
26
+
27
+ import re
28
+ import unicodedata
29
+
30
+ # ---------------------------------------------------------------------------
31
+ # 1. Unicode normalization
32
+ # ---------------------------------------------------------------------------
33
+
34
+
35
+ def nfc(text: str) -> str:
36
+ """Unicode NFC normalize + collapse whitespace + strip."""
37
+ if not isinstance(text, str):
38
+ return ""
39
+ text = unicodedata.normalize("NFC", text)
40
+ text = re.sub(r"\s+", " ", text).strip()
41
+ return text
42
+
43
+
44
+ # ---------------------------------------------------------------------------
45
+ # 2. Language stopwords — common words models hallucinate as names
46
+ # ---------------------------------------------------------------------------
47
+
48
+ LANGUAGE_STOPWORDS: dict[str, set[str]] = {
49
+ "sw": {"jina", "jina langu", "simu", "simu yangu", "barua", "barua pepe", "ninaishi"},
50
+ "vi": {"tôi", "email", "số điện thoại"},
51
+ "tr": {"adım", "e-postam", "telefonum"},
52
+ "id": {"nama", "saya", "email"},
53
+ "pt": {"meu nome"},
54
+ "es": {"me llamo", "mi correo"},
55
+ }
56
+
57
+
58
+ def is_stopword(text: str, language: str | None) -> bool:
59
+ if not language or language not in LANGUAGE_STOPWORDS:
60
+ return False
61
+ return nfc(text).lower() in LANGUAGE_STOPWORDS[language]
62
+
63
+
64
+ def filter_stopwords(entities: list[dict], language: str | None) -> list[dict]:
65
+ return [e for e in entities if not is_stopword(e.get("text", ""), language)]
66
+
67
+
68
+ # ---------------------------------------------------------------------------
69
+ # 3. Same-label overlap deduplication
70
+ # ---------------------------------------------------------------------------
71
+
72
+
73
+ def dedupe_overlapping(entities: list[dict]) -> list[dict]:
74
+ """Drop longer same-label spans that fully contain a shorter same-label span.
75
+
76
+ A clinical downstream prefers specific entities (first_name=An) to loose ones
77
+ (first_name='Nguyễn Văn An'). When the model emits both, we keep the shorter.
78
+ Different-label overlaps are left untouched.
79
+ """
80
+ by_label: dict[str, list[dict]] = {}
81
+ for e in entities:
82
+ by_label.setdefault(e.get("label", ""), []).append(e)
83
+
84
+ kept: list[dict] = []
85
+ for label, group in by_label.items():
86
+ # Sort by length ascending; a span survives only if no shorter same-label
87
+ # span is a substring of it.
88
+ group_sorted = sorted(group, key=lambda x: len(nfc(x.get("text", ""))))
89
+ shorter_texts: list[str] = []
90
+ for e in group_sorted:
91
+ t = nfc(e.get("text", "")).lower()
92
+ if not t:
93
+ continue
94
+ if any(s and s in t and s != t for s in shorter_texts):
95
+ continue # a shorter same-label already covers this
96
+ kept.append(e)
97
+ shorter_texts.append(t)
98
+ return kept
99
+
100
+
101
+ # ---------------------------------------------------------------------------
102
+ # 4. CJK name splitting
103
+ # ---------------------------------------------------------------------------
104
+
105
+ # A small gazetteer of common 2-char Chinese surnames. Extend as needed.
106
+ CHINESE_TWO_CHAR_SURNAMES = {
107
+ "欧阳", "司马", "诸葛", "上官", "夏侯", "东方", "皇甫", "尉迟", "公孙",
108
+ "慕容", "长孙", "宇文", "司徒", "鲜于", "司空", "轩辕", "令狐", "钟离",
109
+ }
110
+
111
+ # Common Japanese surnames (2-char). Tiny set sufficient for the demo; a real
112
+ # system would use a larger dictionary.
113
+ JAPANESE_COMMON_SURNAMES = {
114
+ "佐藤", "鈴木", "高橋", "田中", "伊藤", "渡辺", "山本", "中村", "小林",
115
+ "加藤", "吉田", "山田", "佐々木", "山口", "斎藤", "松本", "井上", "木村",
116
+ "林", "清水",
117
+ }
118
+
119
+ _CJK_RE = re.compile(r"^[\u3400-\u9fff\u3040-\u30ff\uac00-\ud7af]+$")
120
+
121
+
122
+ def _is_cjk(text: str) -> bool:
123
+ return bool(text) and bool(_CJK_RE.match(text))
124
+
125
+
126
+ def _split_korean_name(text: str) -> tuple[str, str] | None:
127
+ # Korean: 1-char surname + 2-char given name is the overwhelming pattern.
128
+ # Only split at 3+ chars; 2-char strings are likely a surname or given alone.
129
+ if len(text) == 3:
130
+ return text[0], text[1:]
131
+ if len(text) == 4:
132
+ return text[:2], text[2:]
133
+ return None
134
+
135
+
136
+ def _split_chinese_name(text: str) -> tuple[str, str] | None:
137
+ # Require 3+ chars. A 2-char Chinese string is almost always a given name
138
+ # on its own (e.g. "小明") rather than a full name to split.
139
+ if len(text) < 3 or len(text) > 4:
140
+ return None
141
+ if text[:2] in CHINESE_TWO_CHAR_SURNAMES:
142
+ return text[:2], text[2:]
143
+ return text[0], text[1:]
144
+
145
+
146
+ def _split_japanese_name(text: str) -> tuple[str, str] | None:
147
+ # Require 3+ chars. A 2-char Japanese string is typically a given name
148
+ # ("太郎", "花子") or a surname alone ("田中", "鈴木") — context-ambiguous,
149
+ # so do nothing. 4-char falls back to 2+2 (typical kanji full name).
150
+ if len(text) < 3:
151
+ return None
152
+ for n in (3, 2):
153
+ if text[:n] in JAPANESE_COMMON_SURNAMES and len(text) > n:
154
+ return text[:n], text[n:]
155
+ if len(text) == 4:
156
+ return text[:2], text[2:]
157
+ return text[:1], text[1:]
158
+
159
+
160
+ def split_cjk_name(text: str, language: str) -> tuple[str, str] | None:
161
+ text = nfc(text)
162
+ if not _is_cjk(text):
163
+ return None
164
+ if language == "ko":
165
+ return _split_korean_name(text)
166
+ if language == "ja":
167
+ return _split_japanese_name(text)
168
+ if language == "zh":
169
+ return _split_chinese_name(text)
170
+ return None
171
+
172
+
173
+ VIETNAMESE_COMMON_SURNAMES = {
174
+ "Nguyễn", "Trần", "Lê", "Phạm", "Hoàng", "Huỳnh", "Phan", "Vũ", "Võ",
175
+ "Đặng", "Bùi", "Đỗ", "Hồ", "Ngô", "Dương", "Lý", "Trịnh", "Đoàn", "Mai",
176
+ }
177
+
178
+
179
+ def _looks_like_vietnamese_surname(text: str) -> bool:
180
+ return nfc(text) in VIETNAMESE_COMMON_SURNAMES
181
+
182
+
183
+ def swap_vietnamese_name_order(entities: list[dict], language: str | None) -> list[dict]:
184
+ """Vietnamese writes names as <family> <middle> <given>. Models trained on
185
+ Western ordering call the first token `first_name` and the last token
186
+ `last_name`, which is the opposite of the Vietnamese convention.
187
+
188
+ We only swap when we can confirm the mistake — specifically, when a value
189
+ labeled `first_name` is a known Vietnamese surname. This avoids breaking
190
+ ground truth that is already labeled correctly.
191
+ """
192
+ if language != "vi":
193
+ return entities
194
+ needs_swap = any(
195
+ e.get("label") == "first_name" and _looks_like_vietnamese_surname(str(e.get("text", "")))
196
+ for e in entities
197
+ )
198
+ if not needs_swap:
199
+ return entities
200
+ swapped: list[dict] = []
201
+ for e in entities:
202
+ lbl = e.get("label")
203
+ if lbl == "first_name":
204
+ swapped.append({**e, "label": "last_name"})
205
+ elif lbl == "last_name":
206
+ swapped.append({**e, "label": "first_name"})
207
+ else:
208
+ swapped.append(e)
209
+ return swapped
210
+
211
+
212
+ def expand_cjk_names(entities: list[dict], language: str | None) -> list[dict]:
213
+ """If a joined CJK name is emitted as first_name / last_name / full_name,
214
+ also emit the split (surname, given_name) pair so matching is generous.
215
+ """
216
+ if language not in {"zh", "ja", "ko"}:
217
+ return entities
218
+ NAME_LABELS = {"first_name", "last_name", "name", "full_name", "person_name"}
219
+ expanded = list(entities)
220
+ seen = {(nfc(e.get("text", "")).lower(), e.get("label", "")) for e in entities}
221
+ for e in entities:
222
+ label = str(e.get("label", "")).lower()
223
+ if label not in NAME_LABELS:
224
+ continue
225
+ text = nfc(e.get("text", ""))
226
+ split = split_cjk_name(text, language)
227
+ if not split:
228
+ continue
229
+ surname, given = split
230
+ for new_text, new_label in [(surname, "last_name"), (given, "first_name")]:
231
+ key = (new_text.lower(), new_label)
232
+ if key not in seen:
233
+ expanded.append({"text": new_text, "label": new_label})
234
+ seen.add(key)
235
+ return expanded
236
+
237
+
238
+ # ---------------------------------------------------------------------------
239
+ # 5. Fuzzy text matching (Slavic case tolerance, substring, NFC)
240
+ # ---------------------------------------------------------------------------
241
+
242
+
243
+ SLAVIC_LANGS = {"ru", "uk", "pl", "cs", "bg", "sk", "sr", "hr"}
244
+
245
+
246
+ def _common_prefix_len(a: str, b: str) -> int:
247
+ n = 0
248
+ for x, y in zip(a, b):
249
+ if x == y:
250
+ n += 1
251
+ else:
252
+ break
253
+ return n
254
+
255
+
256
+ def fuzzy_text_match(a: str, b: str, language: str | None = None) -> bool:
257
+ """Compare two entity text values with production-style tolerance.
258
+
259
+ Returns True if:
260
+ - exact match after NFC + case-fold
261
+ - one is a (word-boundary) substring of the other
262
+ - for Slavic languages, strings share a long common prefix (case inflection)
263
+ """
264
+ a_norm = nfc(a).lower()
265
+ b_norm = nfc(b).lower()
266
+ if not a_norm or not b_norm:
267
+ return False
268
+ if a_norm == b_norm:
269
+ return True
270
+
271
+ # Substring containment (common for "Москве" vs "Москва" isn't substring,
272
+ # but "Seattle, WA" vs "Seattle" is).
273
+ if a_norm in b_norm or b_norm in a_norm:
274
+ # Avoid matching very short substrings inside long ones (e.g. "An" in "Anna").
275
+ shorter, longer = sorted([a_norm, b_norm], key=len)
276
+ if len(shorter) >= 3 and (len(shorter) / len(longer)) >= 0.5:
277
+ return True
278
+
279
+ # Slavic case inflection: Москва / Москве / Москвы share root "Москв"
280
+ if language in SLAVIC_LANGS:
281
+ min_len = min(len(a_norm), len(b_norm))
282
+ cp = _common_prefix_len(a_norm, b_norm)
283
+ if cp >= max(3, min_len - 2):
284
+ return True
285
+
286
+ return False
287
+
288
+
289
+ # ---------------------------------------------------------------------------
290
+ # 6. Top-level postprocess
291
+ # ---------------------------------------------------------------------------
292
+
293
+
294
+ def postprocess_entities(
295
+ entities: list[dict],
296
+ language: str | None = None,
297
+ expand_cjk: bool = True,
298
+ dedupe: bool = True,
299
+ filter_stops: bool = True,
300
+ ) -> list[dict]:
301
+ """Apply the full post-processing pipeline to a list of entity dicts.
302
+
303
+ The order matters: normalize first, then expand CJK splits so both the joined
304
+ and split forms are present, then dedupe same-label overlaps, then filter
305
+ language stopwords.
306
+ """
307
+ if not entities:
308
+ return []
309
+ # Normalize text fields
310
+ normed: list[dict] = []
311
+ for e in entities:
312
+ if not isinstance(e, dict):
313
+ continue
314
+ t = nfc(e.get("text", ""))
315
+ if not t:
316
+ continue
317
+ label = str(e.get("label", "")).strip().lower()
318
+ normed.append({"text": t, "label": label})
319
+
320
+ if expand_cjk:
321
+ normed = expand_cjk_names(normed, language)
322
+ normed = swap_vietnamese_name_order(normed, language)
323
+ if dedupe:
324
+ normed = dedupe_overlapping(normed)
325
+ if filter_stops:
326
+ normed = filter_stopwords(normed, language)
327
+ return normed
328
+
329
+
330
+ def preprocess_text(text: str) -> str:
331
+ """Pre-processing applied before the model sees the input.
332
+
333
+ Mirrors what a clinical pipeline would do to incoming free text:
334
+ - NFC normalize
335
+ - Strip zero-width and control characters
336
+ - Collapse internal whitespace but keep structure (newlines preserved)
337
+ """
338
+ if not isinstance(text, str):
339
+ return ""
340
+ text = unicodedata.normalize("NFC", text)
341
+ # Strip zero-width and bidi control characters that confuse tokenizers.
342
+ text = re.sub(r"[\u200b-\u200f\u202a-\u202e\u2060\ufeff]", "", text)
343
+ # Collapse runs of spaces/tabs but keep newlines.
344
+ text = re.sub(r"[ \t]+", " ", text)
345
+ return text.strip()
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:286acad9b0e27fce778ac429763536accf618ccb6ed72963b6f94685e531c5c7
3
+ size 17077402
tokenizer_config.json ADDED
@@ -0,0 +1,1013 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "bos_token": "<s>",
4
+ "eos_token": "</s>",
5
+ "extra_special_tokens": [
6
+ "<unk>",
7
+ "<s>",
8
+ "</s>",
9
+ "[INST]",
10
+ "[/INST]",
11
+ "[AVAILABLE_TOOLS]",
12
+ "[/AVAILABLE_TOOLS]",
13
+ "[TOOL_RESULTS]",
14
+ "[/TOOL_RESULTS]",
15
+ "[TOOL_CALLS]",
16
+ "[IMG]",
17
+ "<pad>",
18
+ "[IMG_BREAK]",
19
+ "[IMG_END]",
20
+ "[PREFIX]",
21
+ "[MIDDLE]",
22
+ "[SUFFIX]",
23
+ "[SYSTEM_PROMPT]",
24
+ "[/SYSTEM_PROMPT]",
25
+ "[TOOL_CONTENT]",
26
+ "<SPECIAL_20>",
27
+ "<SPECIAL_21>",
28
+ "<SPECIAL_22>",
29
+ "<SPECIAL_23>",
30
+ "[AUDIO]",
31
+ "[BEGIN_AUDIO]",
32
+ "<SPECIAL_26>",
33
+ "<SPECIAL_27>",
34
+ "<SPECIAL_28>",
35
+ "<SPECIAL_29>",
36
+ "<SPECIAL_30>",
37
+ "<SPECIAL_31>",
38
+ "[ARGS]",
39
+ "[CALL_ID]",
40
+ "[THINK]",
41
+ "[/THINK]",
42
+ "<SPECIAL_36>",
43
+ "<SPECIAL_37>",
44
+ "<SPECIAL_38>",
45
+ "<SPECIAL_39>",
46
+ "<SPECIAL_40>",
47
+ "<SPECIAL_41>",
48
+ "<SPECIAL_42>",
49
+ "<SPECIAL_43>",
50
+ "<SPECIAL_44>",
51
+ "<SPECIAL_45>",
52
+ "<SPECIAL_46>",
53
+ "<SPECIAL_47>",
54
+ "<SPECIAL_48>",
55
+ "<SPECIAL_49>",
56
+ "<SPECIAL_50>",
57
+ "<SPECIAL_51>",
58
+ "<SPECIAL_52>",
59
+ "<SPECIAL_53>",
60
+ "<SPECIAL_54>",
61
+ "<SPECIAL_55>",
62
+ "<SPECIAL_56>",
63
+ "<SPECIAL_57>",
64
+ "<SPECIAL_58>",
65
+ "<SPECIAL_59>",
66
+ "<SPECIAL_60>",
67
+ "<SPECIAL_61>",
68
+ "<SPECIAL_62>",
69
+ "<SPECIAL_63>",
70
+ "<SPECIAL_64>",
71
+ "<SPECIAL_65>",
72
+ "<SPECIAL_66>",
73
+ "<SPECIAL_67>",
74
+ "<SPECIAL_68>",
75
+ "<SPECIAL_69>",
76
+ "<SPECIAL_70>",
77
+ "<SPECIAL_71>",
78
+ "<SPECIAL_72>",
79
+ "<SPECIAL_73>",
80
+ "<SPECIAL_74>",
81
+ "<SPECIAL_75>",
82
+ "<SPECIAL_76>",
83
+ "<SPECIAL_77>",
84
+ "<SPECIAL_78>",
85
+ "<SPECIAL_79>",
86
+ "<SPECIAL_80>",
87
+ "<SPECIAL_81>",
88
+ "<SPECIAL_82>",
89
+ "<SPECIAL_83>",
90
+ "<SPECIAL_84>",
91
+ "<SPECIAL_85>",
92
+ "<SPECIAL_86>",
93
+ "<SPECIAL_87>",
94
+ "<SPECIAL_88>",
95
+ "<SPECIAL_89>",
96
+ "<SPECIAL_90>",
97
+ "<SPECIAL_91>",
98
+ "<SPECIAL_92>",
99
+ "<SPECIAL_93>",
100
+ "<SPECIAL_94>",
101
+ "<SPECIAL_95>",
102
+ "<SPECIAL_96>",
103
+ "<SPECIAL_97>",
104
+ "<SPECIAL_98>",
105
+ "<SPECIAL_99>",
106
+ "<SPECIAL_100>",
107
+ "<SPECIAL_101>",
108
+ "<SPECIAL_102>",
109
+ "<SPECIAL_103>",
110
+ "<SPECIAL_104>",
111
+ "<SPECIAL_105>",
112
+ "<SPECIAL_106>",
113
+ "<SPECIAL_107>",
114
+ "<SPECIAL_108>",
115
+ "<SPECIAL_109>",
116
+ "<SPECIAL_110>",
117
+ "<SPECIAL_111>",
118
+ "<SPECIAL_112>",
119
+ "<SPECIAL_113>",
120
+ "<SPECIAL_114>",
121
+ "<SPECIAL_115>",
122
+ "<SPECIAL_116>",
123
+ "<SPECIAL_117>",
124
+ "<SPECIAL_118>",
125
+ "<SPECIAL_119>",
126
+ "<SPECIAL_120>",
127
+ "<SPECIAL_121>",
128
+ "<SPECIAL_122>",
129
+ "<SPECIAL_123>",
130
+ "<SPECIAL_124>",
131
+ "<SPECIAL_125>",
132
+ "<SPECIAL_126>",
133
+ "<SPECIAL_127>",
134
+ "<SPECIAL_128>",
135
+ "<SPECIAL_129>",
136
+ "<SPECIAL_130>",
137
+ "<SPECIAL_131>",
138
+ "<SPECIAL_132>",
139
+ "<SPECIAL_133>",
140
+ "<SPECIAL_134>",
141
+ "<SPECIAL_135>",
142
+ "<SPECIAL_136>",
143
+ "<SPECIAL_137>",
144
+ "<SPECIAL_138>",
145
+ "<SPECIAL_139>",
146
+ "<SPECIAL_140>",
147
+ "<SPECIAL_141>",
148
+ "<SPECIAL_142>",
149
+ "<SPECIAL_143>",
150
+ "<SPECIAL_144>",
151
+ "<SPECIAL_145>",
152
+ "<SPECIAL_146>",
153
+ "<SPECIAL_147>",
154
+ "<SPECIAL_148>",
155
+ "<SPECIAL_149>",
156
+ "<SPECIAL_150>",
157
+ "<SPECIAL_151>",
158
+ "<SPECIAL_152>",
159
+ "<SPECIAL_153>",
160
+ "<SPECIAL_154>",
161
+ "<SPECIAL_155>",
162
+ "<SPECIAL_156>",
163
+ "<SPECIAL_157>",
164
+ "<SPECIAL_158>",
165
+ "<SPECIAL_159>",
166
+ "<SPECIAL_160>",
167
+ "<SPECIAL_161>",
168
+ "<SPECIAL_162>",
169
+ "<SPECIAL_163>",
170
+ "<SPECIAL_164>",
171
+ "<SPECIAL_165>",
172
+ "<SPECIAL_166>",
173
+ "<SPECIAL_167>",
174
+ "<SPECIAL_168>",
175
+ "<SPECIAL_169>",
176
+ "<SPECIAL_170>",
177
+ "<SPECIAL_171>",
178
+ "<SPECIAL_172>",
179
+ "<SPECIAL_173>",
180
+ "<SPECIAL_174>",
181
+ "<SPECIAL_175>",
182
+ "<SPECIAL_176>",
183
+ "<SPECIAL_177>",
184
+ "<SPECIAL_178>",
185
+ "<SPECIAL_179>",
186
+ "<SPECIAL_180>",
187
+ "<SPECIAL_181>",
188
+ "<SPECIAL_182>",
189
+ "<SPECIAL_183>",
190
+ "<SPECIAL_184>",
191
+ "<SPECIAL_185>",
192
+ "<SPECIAL_186>",
193
+ "<SPECIAL_187>",
194
+ "<SPECIAL_188>",
195
+ "<SPECIAL_189>",
196
+ "<SPECIAL_190>",
197
+ "<SPECIAL_191>",
198
+ "<SPECIAL_192>",
199
+ "<SPECIAL_193>",
200
+ "<SPECIAL_194>",
201
+ "<SPECIAL_195>",
202
+ "<SPECIAL_196>",
203
+ "<SPECIAL_197>",
204
+ "<SPECIAL_198>",
205
+ "<SPECIAL_199>",
206
+ "<SPECIAL_200>",
207
+ "<SPECIAL_201>",
208
+ "<SPECIAL_202>",
209
+ "<SPECIAL_203>",
210
+ "<SPECIAL_204>",
211
+ "<SPECIAL_205>",
212
+ "<SPECIAL_206>",
213
+ "<SPECIAL_207>",
214
+ "<SPECIAL_208>",
215
+ "<SPECIAL_209>",
216
+ "<SPECIAL_210>",
217
+ "<SPECIAL_211>",
218
+ "<SPECIAL_212>",
219
+ "<SPECIAL_213>",
220
+ "<SPECIAL_214>",
221
+ "<SPECIAL_215>",
222
+ "<SPECIAL_216>",
223
+ "<SPECIAL_217>",
224
+ "<SPECIAL_218>",
225
+ "<SPECIAL_219>",
226
+ "<SPECIAL_220>",
227
+ "<SPECIAL_221>",
228
+ "<SPECIAL_222>",
229
+ "<SPECIAL_223>",
230
+ "<SPECIAL_224>",
231
+ "<SPECIAL_225>",
232
+ "<SPECIAL_226>",
233
+ "<SPECIAL_227>",
234
+ "<SPECIAL_228>",
235
+ "<SPECIAL_229>",
236
+ "<SPECIAL_230>",
237
+ "<SPECIAL_231>",
238
+ "<SPECIAL_232>",
239
+ "<SPECIAL_233>",
240
+ "<SPECIAL_234>",
241
+ "<SPECIAL_235>",
242
+ "<SPECIAL_236>",
243
+ "<SPECIAL_237>",
244
+ "<SPECIAL_238>",
245
+ "<SPECIAL_239>",
246
+ "<SPECIAL_240>",
247
+ "<SPECIAL_241>",
248
+ "<SPECIAL_242>",
249
+ "<SPECIAL_243>",
250
+ "<SPECIAL_244>",
251
+ "<SPECIAL_245>",
252
+ "<SPECIAL_246>",
253
+ "<SPECIAL_247>",
254
+ "<SPECIAL_248>",
255
+ "<SPECIAL_249>",
256
+ "<SPECIAL_250>",
257
+ "<SPECIAL_251>",
258
+ "<SPECIAL_252>",
259
+ "<SPECIAL_253>",
260
+ "<SPECIAL_254>",
261
+ "<SPECIAL_255>",
262
+ "<SPECIAL_256>",
263
+ "<SPECIAL_257>",
264
+ "<SPECIAL_258>",
265
+ "<SPECIAL_259>",
266
+ "<SPECIAL_260>",
267
+ "<SPECIAL_261>",
268
+ "<SPECIAL_262>",
269
+ "<SPECIAL_263>",
270
+ "<SPECIAL_264>",
271
+ "<SPECIAL_265>",
272
+ "<SPECIAL_266>",
273
+ "<SPECIAL_267>",
274
+ "<SPECIAL_268>",
275
+ "<SPECIAL_269>",
276
+ "<SPECIAL_270>",
277
+ "<SPECIAL_271>",
278
+ "<SPECIAL_272>",
279
+ "<SPECIAL_273>",
280
+ "<SPECIAL_274>",
281
+ "<SPECIAL_275>",
282
+ "<SPECIAL_276>",
283
+ "<SPECIAL_277>",
284
+ "<SPECIAL_278>",
285
+ "<SPECIAL_279>",
286
+ "<SPECIAL_280>",
287
+ "<SPECIAL_281>",
288
+ "<SPECIAL_282>",
289
+ "<SPECIAL_283>",
290
+ "<SPECIAL_284>",
291
+ "<SPECIAL_285>",
292
+ "<SPECIAL_286>",
293
+ "<SPECIAL_287>",
294
+ "<SPECIAL_288>",
295
+ "<SPECIAL_289>",
296
+ "<SPECIAL_290>",
297
+ "<SPECIAL_291>",
298
+ "<SPECIAL_292>",
299
+ "<SPECIAL_293>",
300
+ "<SPECIAL_294>",
301
+ "<SPECIAL_295>",
302
+ "<SPECIAL_296>",
303
+ "<SPECIAL_297>",
304
+ "<SPECIAL_298>",
305
+ "<SPECIAL_299>",
306
+ "<SPECIAL_300>",
307
+ "<SPECIAL_301>",
308
+ "<SPECIAL_302>",
309
+ "<SPECIAL_303>",
310
+ "<SPECIAL_304>",
311
+ "<SPECIAL_305>",
312
+ "<SPECIAL_306>",
313
+ "<SPECIAL_307>",
314
+ "<SPECIAL_308>",
315
+ "<SPECIAL_309>",
316
+ "<SPECIAL_310>",
317
+ "<SPECIAL_311>",
318
+ "<SPECIAL_312>",
319
+ "<SPECIAL_313>",
320
+ "<SPECIAL_314>",
321
+ "<SPECIAL_315>",
322
+ "<SPECIAL_316>",
323
+ "<SPECIAL_317>",
324
+ "<SPECIAL_318>",
325
+ "<SPECIAL_319>",
326
+ "<SPECIAL_320>",
327
+ "<SPECIAL_321>",
328
+ "<SPECIAL_322>",
329
+ "<SPECIAL_323>",
330
+ "<SPECIAL_324>",
331
+ "<SPECIAL_325>",
332
+ "<SPECIAL_326>",
333
+ "<SPECIAL_327>",
334
+ "<SPECIAL_328>",
335
+ "<SPECIAL_329>",
336
+ "<SPECIAL_330>",
337
+ "<SPECIAL_331>",
338
+ "<SPECIAL_332>",
339
+ "<SPECIAL_333>",
340
+ "<SPECIAL_334>",
341
+ "<SPECIAL_335>",
342
+ "<SPECIAL_336>",
343
+ "<SPECIAL_337>",
344
+ "<SPECIAL_338>",
345
+ "<SPECIAL_339>",
346
+ "<SPECIAL_340>",
347
+ "<SPECIAL_341>",
348
+ "<SPECIAL_342>",
349
+ "<SPECIAL_343>",
350
+ "<SPECIAL_344>",
351
+ "<SPECIAL_345>",
352
+ "<SPECIAL_346>",
353
+ "<SPECIAL_347>",
354
+ "<SPECIAL_348>",
355
+ "<SPECIAL_349>",
356
+ "<SPECIAL_350>",
357
+ "<SPECIAL_351>",
358
+ "<SPECIAL_352>",
359
+ "<SPECIAL_353>",
360
+ "<SPECIAL_354>",
361
+ "<SPECIAL_355>",
362
+ "<SPECIAL_356>",
363
+ "<SPECIAL_357>",
364
+ "<SPECIAL_358>",
365
+ "<SPECIAL_359>",
366
+ "<SPECIAL_360>",
367
+ "<SPECIAL_361>",
368
+ "<SPECIAL_362>",
369
+ "<SPECIAL_363>",
370
+ "<SPECIAL_364>",
371
+ "<SPECIAL_365>",
372
+ "<SPECIAL_366>",
373
+ "<SPECIAL_367>",
374
+ "<SPECIAL_368>",
375
+ "<SPECIAL_369>",
376
+ "<SPECIAL_370>",
377
+ "<SPECIAL_371>",
378
+ "<SPECIAL_372>",
379
+ "<SPECIAL_373>",
380
+ "<SPECIAL_374>",
381
+ "<SPECIAL_375>",
382
+ "<SPECIAL_376>",
383
+ "<SPECIAL_377>",
384
+ "<SPECIAL_378>",
385
+ "<SPECIAL_379>",
386
+ "<SPECIAL_380>",
387
+ "<SPECIAL_381>",
388
+ "<SPECIAL_382>",
389
+ "<SPECIAL_383>",
390
+ "<SPECIAL_384>",
391
+ "<SPECIAL_385>",
392
+ "<SPECIAL_386>",
393
+ "<SPECIAL_387>",
394
+ "<SPECIAL_388>",
395
+ "<SPECIAL_389>",
396
+ "<SPECIAL_390>",
397
+ "<SPECIAL_391>",
398
+ "<SPECIAL_392>",
399
+ "<SPECIAL_393>",
400
+ "<SPECIAL_394>",
401
+ "<SPECIAL_395>",
402
+ "<SPECIAL_396>",
403
+ "<SPECIAL_397>",
404
+ "<SPECIAL_398>",
405
+ "<SPECIAL_399>",
406
+ "<SPECIAL_400>",
407
+ "<SPECIAL_401>",
408
+ "<SPECIAL_402>",
409
+ "<SPECIAL_403>",
410
+ "<SPECIAL_404>",
411
+ "<SPECIAL_405>",
412
+ "<SPECIAL_406>",
413
+ "<SPECIAL_407>",
414
+ "<SPECIAL_408>",
415
+ "<SPECIAL_409>",
416
+ "<SPECIAL_410>",
417
+ "<SPECIAL_411>",
418
+ "<SPECIAL_412>",
419
+ "<SPECIAL_413>",
420
+ "<SPECIAL_414>",
421
+ "<SPECIAL_415>",
422
+ "<SPECIAL_416>",
423
+ "<SPECIAL_417>",
424
+ "<SPECIAL_418>",
425
+ "<SPECIAL_419>",
426
+ "<SPECIAL_420>",
427
+ "<SPECIAL_421>",
428
+ "<SPECIAL_422>",
429
+ "<SPECIAL_423>",
430
+ "<SPECIAL_424>",
431
+ "<SPECIAL_425>",
432
+ "<SPECIAL_426>",
433
+ "<SPECIAL_427>",
434
+ "<SPECIAL_428>",
435
+ "<SPECIAL_429>",
436
+ "<SPECIAL_430>",
437
+ "<SPECIAL_431>",
438
+ "<SPECIAL_432>",
439
+ "<SPECIAL_433>",
440
+ "<SPECIAL_434>",
441
+ "<SPECIAL_435>",
442
+ "<SPECIAL_436>",
443
+ "<SPECIAL_437>",
444
+ "<SPECIAL_438>",
445
+ "<SPECIAL_439>",
446
+ "<SPECIAL_440>",
447
+ "<SPECIAL_441>",
448
+ "<SPECIAL_442>",
449
+ "<SPECIAL_443>",
450
+ "<SPECIAL_444>",
451
+ "<SPECIAL_445>",
452
+ "<SPECIAL_446>",
453
+ "<SPECIAL_447>",
454
+ "<SPECIAL_448>",
455
+ "<SPECIAL_449>",
456
+ "<SPECIAL_450>",
457
+ "<SPECIAL_451>",
458
+ "<SPECIAL_452>",
459
+ "<SPECIAL_453>",
460
+ "<SPECIAL_454>",
461
+ "<SPECIAL_455>",
462
+ "<SPECIAL_456>",
463
+ "<SPECIAL_457>",
464
+ "<SPECIAL_458>",
465
+ "<SPECIAL_459>",
466
+ "<SPECIAL_460>",
467
+ "<SPECIAL_461>",
468
+ "<SPECIAL_462>",
469
+ "<SPECIAL_463>",
470
+ "<SPECIAL_464>",
471
+ "<SPECIAL_465>",
472
+ "<SPECIAL_466>",
473
+ "<SPECIAL_467>",
474
+ "<SPECIAL_468>",
475
+ "<SPECIAL_469>",
476
+ "<SPECIAL_470>",
477
+ "<SPECIAL_471>",
478
+ "<SPECIAL_472>",
479
+ "<SPECIAL_473>",
480
+ "<SPECIAL_474>",
481
+ "<SPECIAL_475>",
482
+ "<SPECIAL_476>",
483
+ "<SPECIAL_477>",
484
+ "<SPECIAL_478>",
485
+ "<SPECIAL_479>",
486
+ "<SPECIAL_480>",
487
+ "<SPECIAL_481>",
488
+ "<SPECIAL_482>",
489
+ "<SPECIAL_483>",
490
+ "<SPECIAL_484>",
491
+ "<SPECIAL_485>",
492
+ "<SPECIAL_486>",
493
+ "<SPECIAL_487>",
494
+ "<SPECIAL_488>",
495
+ "<SPECIAL_489>",
496
+ "<SPECIAL_490>",
497
+ "<SPECIAL_491>",
498
+ "<SPECIAL_492>",
499
+ "<SPECIAL_493>",
500
+ "<SPECIAL_494>",
501
+ "<SPECIAL_495>",
502
+ "<SPECIAL_496>",
503
+ "<SPECIAL_497>",
504
+ "<SPECIAL_498>",
505
+ "<SPECIAL_499>",
506
+ "<SPECIAL_500>",
507
+ "<SPECIAL_501>",
508
+ "<SPECIAL_502>",
509
+ "<SPECIAL_503>",
510
+ "<SPECIAL_504>",
511
+ "<SPECIAL_505>",
512
+ "<SPECIAL_506>",
513
+ "<SPECIAL_507>",
514
+ "<SPECIAL_508>",
515
+ "<SPECIAL_509>",
516
+ "<SPECIAL_510>",
517
+ "<SPECIAL_511>",
518
+ "<SPECIAL_512>",
519
+ "<SPECIAL_513>",
520
+ "<SPECIAL_514>",
521
+ "<SPECIAL_515>",
522
+ "<SPECIAL_516>",
523
+ "<SPECIAL_517>",
524
+ "<SPECIAL_518>",
525
+ "<SPECIAL_519>",
526
+ "<SPECIAL_520>",
527
+ "<SPECIAL_521>",
528
+ "<SPECIAL_522>",
529
+ "<SPECIAL_523>",
530
+ "<SPECIAL_524>",
531
+ "<SPECIAL_525>",
532
+ "<SPECIAL_526>",
533
+ "<SPECIAL_527>",
534
+ "<SPECIAL_528>",
535
+ "<SPECIAL_529>",
536
+ "<SPECIAL_530>",
537
+ "<SPECIAL_531>",
538
+ "<SPECIAL_532>",
539
+ "<SPECIAL_533>",
540
+ "<SPECIAL_534>",
541
+ "<SPECIAL_535>",
542
+ "<SPECIAL_536>",
543
+ "<SPECIAL_537>",
544
+ "<SPECIAL_538>",
545
+ "<SPECIAL_539>",
546
+ "<SPECIAL_540>",
547
+ "<SPECIAL_541>",
548
+ "<SPECIAL_542>",
549
+ "<SPECIAL_543>",
550
+ "<SPECIAL_544>",
551
+ "<SPECIAL_545>",
552
+ "<SPECIAL_546>",
553
+ "<SPECIAL_547>",
554
+ "<SPECIAL_548>",
555
+ "<SPECIAL_549>",
556
+ "<SPECIAL_550>",
557
+ "<SPECIAL_551>",
558
+ "<SPECIAL_552>",
559
+ "<SPECIAL_553>",
560
+ "<SPECIAL_554>",
561
+ "<SPECIAL_555>",
562
+ "<SPECIAL_556>",
563
+ "<SPECIAL_557>",
564
+ "<SPECIAL_558>",
565
+ "<SPECIAL_559>",
566
+ "<SPECIAL_560>",
567
+ "<SPECIAL_561>",
568
+ "<SPECIAL_562>",
569
+ "<SPECIAL_563>",
570
+ "<SPECIAL_564>",
571
+ "<SPECIAL_565>",
572
+ "<SPECIAL_566>",
573
+ "<SPECIAL_567>",
574
+ "<SPECIAL_568>",
575
+ "<SPECIAL_569>",
576
+ "<SPECIAL_570>",
577
+ "<SPECIAL_571>",
578
+ "<SPECIAL_572>",
579
+ "<SPECIAL_573>",
580
+ "<SPECIAL_574>",
581
+ "<SPECIAL_575>",
582
+ "<SPECIAL_576>",
583
+ "<SPECIAL_577>",
584
+ "<SPECIAL_578>",
585
+ "<SPECIAL_579>",
586
+ "<SPECIAL_580>",
587
+ "<SPECIAL_581>",
588
+ "<SPECIAL_582>",
589
+ "<SPECIAL_583>",
590
+ "<SPECIAL_584>",
591
+ "<SPECIAL_585>",
592
+ "<SPECIAL_586>",
593
+ "<SPECIAL_587>",
594
+ "<SPECIAL_588>",
595
+ "<SPECIAL_589>",
596
+ "<SPECIAL_590>",
597
+ "<SPECIAL_591>",
598
+ "<SPECIAL_592>",
599
+ "<SPECIAL_593>",
600
+ "<SPECIAL_594>",
601
+ "<SPECIAL_595>",
602
+ "<SPECIAL_596>",
603
+ "<SPECIAL_597>",
604
+ "<SPECIAL_598>",
605
+ "<SPECIAL_599>",
606
+ "<SPECIAL_600>",
607
+ "<SPECIAL_601>",
608
+ "<SPECIAL_602>",
609
+ "<SPECIAL_603>",
610
+ "<SPECIAL_604>",
611
+ "<SPECIAL_605>",
612
+ "<SPECIAL_606>",
613
+ "<SPECIAL_607>",
614
+ "<SPECIAL_608>",
615
+ "<SPECIAL_609>",
616
+ "<SPECIAL_610>",
617
+ "<SPECIAL_611>",
618
+ "<SPECIAL_612>",
619
+ "<SPECIAL_613>",
620
+ "<SPECIAL_614>",
621
+ "<SPECIAL_615>",
622
+ "<SPECIAL_616>",
623
+ "<SPECIAL_617>",
624
+ "<SPECIAL_618>",
625
+ "<SPECIAL_619>",
626
+ "<SPECIAL_620>",
627
+ "<SPECIAL_621>",
628
+ "<SPECIAL_622>",
629
+ "<SPECIAL_623>",
630
+ "<SPECIAL_624>",
631
+ "<SPECIAL_625>",
632
+ "<SPECIAL_626>",
633
+ "<SPECIAL_627>",
634
+ "<SPECIAL_628>",
635
+ "<SPECIAL_629>",
636
+ "<SPECIAL_630>",
637
+ "<SPECIAL_631>",
638
+ "<SPECIAL_632>",
639
+ "<SPECIAL_633>",
640
+ "<SPECIAL_634>",
641
+ "<SPECIAL_635>",
642
+ "<SPECIAL_636>",
643
+ "<SPECIAL_637>",
644
+ "<SPECIAL_638>",
645
+ "<SPECIAL_639>",
646
+ "<SPECIAL_640>",
647
+ "<SPECIAL_641>",
648
+ "<SPECIAL_642>",
649
+ "<SPECIAL_643>",
650
+ "<SPECIAL_644>",
651
+ "<SPECIAL_645>",
652
+ "<SPECIAL_646>",
653
+ "<SPECIAL_647>",
654
+ "<SPECIAL_648>",
655
+ "<SPECIAL_649>",
656
+ "<SPECIAL_650>",
657
+ "<SPECIAL_651>",
658
+ "<SPECIAL_652>",
659
+ "<SPECIAL_653>",
660
+ "<SPECIAL_654>",
661
+ "<SPECIAL_655>",
662
+ "<SPECIAL_656>",
663
+ "<SPECIAL_657>",
664
+ "<SPECIAL_658>",
665
+ "<SPECIAL_659>",
666
+ "<SPECIAL_660>",
667
+ "<SPECIAL_661>",
668
+ "<SPECIAL_662>",
669
+ "<SPECIAL_663>",
670
+ "<SPECIAL_664>",
671
+ "<SPECIAL_665>",
672
+ "<SPECIAL_666>",
673
+ "<SPECIAL_667>",
674
+ "<SPECIAL_668>",
675
+ "<SPECIAL_669>",
676
+ "<SPECIAL_670>",
677
+ "<SPECIAL_671>",
678
+ "<SPECIAL_672>",
679
+ "<SPECIAL_673>",
680
+ "<SPECIAL_674>",
681
+ "<SPECIAL_675>",
682
+ "<SPECIAL_676>",
683
+ "<SPECIAL_677>",
684
+ "<SPECIAL_678>",
685
+ "<SPECIAL_679>",
686
+ "<SPECIAL_680>",
687
+ "<SPECIAL_681>",
688
+ "<SPECIAL_682>",
689
+ "<SPECIAL_683>",
690
+ "<SPECIAL_684>",
691
+ "<SPECIAL_685>",
692
+ "<SPECIAL_686>",
693
+ "<SPECIAL_687>",
694
+ "<SPECIAL_688>",
695
+ "<SPECIAL_689>",
696
+ "<SPECIAL_690>",
697
+ "<SPECIAL_691>",
698
+ "<SPECIAL_692>",
699
+ "<SPECIAL_693>",
700
+ "<SPECIAL_694>",
701
+ "<SPECIAL_695>",
702
+ "<SPECIAL_696>",
703
+ "<SPECIAL_697>",
704
+ "<SPECIAL_698>",
705
+ "<SPECIAL_699>",
706
+ "<SPECIAL_700>",
707
+ "<SPECIAL_701>",
708
+ "<SPECIAL_702>",
709
+ "<SPECIAL_703>",
710
+ "<SPECIAL_704>",
711
+ "<SPECIAL_705>",
712
+ "<SPECIAL_706>",
713
+ "<SPECIAL_707>",
714
+ "<SPECIAL_708>",
715
+ "<SPECIAL_709>",
716
+ "<SPECIAL_710>",
717
+ "<SPECIAL_711>",
718
+ "<SPECIAL_712>",
719
+ "<SPECIAL_713>",
720
+ "<SPECIAL_714>",
721
+ "<SPECIAL_715>",
722
+ "<SPECIAL_716>",
723
+ "<SPECIAL_717>",
724
+ "<SPECIAL_718>",
725
+ "<SPECIAL_719>",
726
+ "<SPECIAL_720>",
727
+ "<SPECIAL_721>",
728
+ "<SPECIAL_722>",
729
+ "<SPECIAL_723>",
730
+ "<SPECIAL_724>",
731
+ "<SPECIAL_725>",
732
+ "<SPECIAL_726>",
733
+ "<SPECIAL_727>",
734
+ "<SPECIAL_728>",
735
+ "<SPECIAL_729>",
736
+ "<SPECIAL_730>",
737
+ "<SPECIAL_731>",
738
+ "<SPECIAL_732>",
739
+ "<SPECIAL_733>",
740
+ "<SPECIAL_734>",
741
+ "<SPECIAL_735>",
742
+ "<SPECIAL_736>",
743
+ "<SPECIAL_737>",
744
+ "<SPECIAL_738>",
745
+ "<SPECIAL_739>",
746
+ "<SPECIAL_740>",
747
+ "<SPECIAL_741>",
748
+ "<SPECIAL_742>",
749
+ "<SPECIAL_743>",
750
+ "<SPECIAL_744>",
751
+ "<SPECIAL_745>",
752
+ "<SPECIAL_746>",
753
+ "<SPECIAL_747>",
754
+ "<SPECIAL_748>",
755
+ "<SPECIAL_749>",
756
+ "<SPECIAL_750>",
757
+ "<SPECIAL_751>",
758
+ "<SPECIAL_752>",
759
+ "<SPECIAL_753>",
760
+ "<SPECIAL_754>",
761
+ "<SPECIAL_755>",
762
+ "<SPECIAL_756>",
763
+ "<SPECIAL_757>",
764
+ "<SPECIAL_758>",
765
+ "<SPECIAL_759>",
766
+ "<SPECIAL_760>",
767
+ "<SPECIAL_761>",
768
+ "<SPECIAL_762>",
769
+ "<SPECIAL_763>",
770
+ "<SPECIAL_764>",
771
+ "<SPECIAL_765>",
772
+ "<SPECIAL_766>",
773
+ "<SPECIAL_767>",
774
+ "<SPECIAL_768>",
775
+ "<SPECIAL_769>",
776
+ "<SPECIAL_770>",
777
+ "<SPECIAL_771>",
778
+ "<SPECIAL_772>",
779
+ "<SPECIAL_773>",
780
+ "<SPECIAL_774>",
781
+ "<SPECIAL_775>",
782
+ "<SPECIAL_776>",
783
+ "<SPECIAL_777>",
784
+ "<SPECIAL_778>",
785
+ "<SPECIAL_779>",
786
+ "<SPECIAL_780>",
787
+ "<SPECIAL_781>",
788
+ "<SPECIAL_782>",
789
+ "<SPECIAL_783>",
790
+ "<SPECIAL_784>",
791
+ "<SPECIAL_785>",
792
+ "<SPECIAL_786>",
793
+ "<SPECIAL_787>",
794
+ "<SPECIAL_788>",
795
+ "<SPECIAL_789>",
796
+ "<SPECIAL_790>",
797
+ "<SPECIAL_791>",
798
+ "<SPECIAL_792>",
799
+ "<SPECIAL_793>",
800
+ "<SPECIAL_794>",
801
+ "<SPECIAL_795>",
802
+ "<SPECIAL_796>",
803
+ "<SPECIAL_797>",
804
+ "<SPECIAL_798>",
805
+ "<SPECIAL_799>",
806
+ "<SPECIAL_800>",
807
+ "<SPECIAL_801>",
808
+ "<SPECIAL_802>",
809
+ "<SPECIAL_803>",
810
+ "<SPECIAL_804>",
811
+ "<SPECIAL_805>",
812
+ "<SPECIAL_806>",
813
+ "<SPECIAL_807>",
814
+ "<SPECIAL_808>",
815
+ "<SPECIAL_809>",
816
+ "<SPECIAL_810>",
817
+ "<SPECIAL_811>",
818
+ "<SPECIAL_812>",
819
+ "<SPECIAL_813>",
820
+ "<SPECIAL_814>",
821
+ "<SPECIAL_815>",
822
+ "<SPECIAL_816>",
823
+ "<SPECIAL_817>",
824
+ "<SPECIAL_818>",
825
+ "<SPECIAL_819>",
826
+ "<SPECIAL_820>",
827
+ "<SPECIAL_821>",
828
+ "<SPECIAL_822>",
829
+ "<SPECIAL_823>",
830
+ "<SPECIAL_824>",
831
+ "<SPECIAL_825>",
832
+ "<SPECIAL_826>",
833
+ "<SPECIAL_827>",
834
+ "<SPECIAL_828>",
835
+ "<SPECIAL_829>",
836
+ "<SPECIAL_830>",
837
+ "<SPECIAL_831>",
838
+ "<SPECIAL_832>",
839
+ "<SPECIAL_833>",
840
+ "<SPECIAL_834>",
841
+ "<SPECIAL_835>",
842
+ "<SPECIAL_836>",
843
+ "<SPECIAL_837>",
844
+ "<SPECIAL_838>",
845
+ "<SPECIAL_839>",
846
+ "<SPECIAL_840>",
847
+ "<SPECIAL_841>",
848
+ "<SPECIAL_842>",
849
+ "<SPECIAL_843>",
850
+ "<SPECIAL_844>",
851
+ "<SPECIAL_845>",
852
+ "<SPECIAL_846>",
853
+ "<SPECIAL_847>",
854
+ "<SPECIAL_848>",
855
+ "<SPECIAL_849>",
856
+ "<SPECIAL_850>",
857
+ "<SPECIAL_851>",
858
+ "<SPECIAL_852>",
859
+ "<SPECIAL_853>",
860
+ "<SPECIAL_854>",
861
+ "<SPECIAL_855>",
862
+ "<SPECIAL_856>",
863
+ "<SPECIAL_857>",
864
+ "<SPECIAL_858>",
865
+ "<SPECIAL_859>",
866
+ "<SPECIAL_860>",
867
+ "<SPECIAL_861>",
868
+ "<SPECIAL_862>",
869
+ "<SPECIAL_863>",
870
+ "<SPECIAL_864>",
871
+ "<SPECIAL_865>",
872
+ "<SPECIAL_866>",
873
+ "<SPECIAL_867>",
874
+ "<SPECIAL_868>",
875
+ "<SPECIAL_869>",
876
+ "<SPECIAL_870>",
877
+ "<SPECIAL_871>",
878
+ "<SPECIAL_872>",
879
+ "<SPECIAL_873>",
880
+ "<SPECIAL_874>",
881
+ "<SPECIAL_875>",
882
+ "<SPECIAL_876>",
883
+ "<SPECIAL_877>",
884
+ "<SPECIAL_878>",
885
+ "<SPECIAL_879>",
886
+ "<SPECIAL_880>",
887
+ "<SPECIAL_881>",
888
+ "<SPECIAL_882>",
889
+ "<SPECIAL_883>",
890
+ "<SPECIAL_884>",
891
+ "<SPECIAL_885>",
892
+ "<SPECIAL_886>",
893
+ "<SPECIAL_887>",
894
+ "<SPECIAL_888>",
895
+ "<SPECIAL_889>",
896
+ "<SPECIAL_890>",
897
+ "<SPECIAL_891>",
898
+ "<SPECIAL_892>",
899
+ "<SPECIAL_893>",
900
+ "<SPECIAL_894>",
901
+ "<SPECIAL_895>",
902
+ "<SPECIAL_896>",
903
+ "<SPECIAL_897>",
904
+ "<SPECIAL_898>",
905
+ "<SPECIAL_899>",
906
+ "<SPECIAL_900>",
907
+ "<SPECIAL_901>",
908
+ "<SPECIAL_902>",
909
+ "<SPECIAL_903>",
910
+ "<SPECIAL_904>",
911
+ "<SPECIAL_905>",
912
+ "<SPECIAL_906>",
913
+ "<SPECIAL_907>",
914
+ "<SPECIAL_908>",
915
+ "<SPECIAL_909>",
916
+ "<SPECIAL_910>",
917
+ "<SPECIAL_911>",
918
+ "<SPECIAL_912>",
919
+ "<SPECIAL_913>",
920
+ "<SPECIAL_914>",
921
+ "<SPECIAL_915>",
922
+ "<SPECIAL_916>",
923
+ "<SPECIAL_917>",
924
+ "<SPECIAL_918>",
925
+ "<SPECIAL_919>",
926
+ "<SPECIAL_920>",
927
+ "<SPECIAL_921>",
928
+ "<SPECIAL_922>",
929
+ "<SPECIAL_923>",
930
+ "<SPECIAL_924>",
931
+ "<SPECIAL_925>",
932
+ "<SPECIAL_926>",
933
+ "<SPECIAL_927>",
934
+ "<SPECIAL_928>",
935
+ "<SPECIAL_929>",
936
+ "<SPECIAL_930>",
937
+ "<SPECIAL_931>",
938
+ "<SPECIAL_932>",
939
+ "<SPECIAL_933>",
940
+ "<SPECIAL_934>",
941
+ "<SPECIAL_935>",
942
+ "<SPECIAL_936>",
943
+ "<SPECIAL_937>",
944
+ "<SPECIAL_938>",
945
+ "<SPECIAL_939>",
946
+ "<SPECIAL_940>",
947
+ "<SPECIAL_941>",
948
+ "<SPECIAL_942>",
949
+ "<SPECIAL_943>",
950
+ "<SPECIAL_944>",
951
+ "<SPECIAL_945>",
952
+ "<SPECIAL_946>",
953
+ "<SPECIAL_947>",
954
+ "<SPECIAL_948>",
955
+ "<SPECIAL_949>",
956
+ "<SPECIAL_950>",
957
+ "<SPECIAL_951>",
958
+ "<SPECIAL_952>",
959
+ "<SPECIAL_953>",
960
+ "<SPECIAL_954>",
961
+ "<SPECIAL_955>",
962
+ "<SPECIAL_956>",
963
+ "<SPECIAL_957>",
964
+ "<SPECIAL_958>",
965
+ "<SPECIAL_959>",
966
+ "<SPECIAL_960>",
967
+ "<SPECIAL_961>",
968
+ "<SPECIAL_962>",
969
+ "<SPECIAL_963>",
970
+ "<SPECIAL_964>",
971
+ "<SPECIAL_965>",
972
+ "<SPECIAL_966>",
973
+ "<SPECIAL_967>",
974
+ "<SPECIAL_968>",
975
+ "<SPECIAL_969>",
976
+ "<SPECIAL_970>",
977
+ "<SPECIAL_971>",
978
+ "<SPECIAL_972>",
979
+ "<SPECIAL_973>",
980
+ "<SPECIAL_974>",
981
+ "<SPECIAL_975>",
982
+ "<SPECIAL_976>",
983
+ "<SPECIAL_977>",
984
+ "<SPECIAL_978>",
985
+ "<SPECIAL_979>",
986
+ "<SPECIAL_980>",
987
+ "<SPECIAL_981>",
988
+ "<SPECIAL_982>",
989
+ "<SPECIAL_983>",
990
+ "<SPECIAL_984>",
991
+ "<SPECIAL_985>",
992
+ "<SPECIAL_986>",
993
+ "<SPECIAL_987>",
994
+ "<SPECIAL_988>",
995
+ "<SPECIAL_989>",
996
+ "<SPECIAL_990>",
997
+ "<SPECIAL_991>",
998
+ "<SPECIAL_992>",
999
+ "<SPECIAL_993>",
1000
+ "<SPECIAL_994>",
1001
+ "<SPECIAL_995>",
1002
+ "<SPECIAL_996>",
1003
+ "<SPECIAL_997>",
1004
+ "<SPECIAL_998>",
1005
+ "<SPECIAL_999>"
1006
+ ],
1007
+ "is_local": false,
1008
+ "model_max_length": 1000000000000000019884624838656,
1009
+ "pad_token": "<pad>",
1010
+ "processor_class": "PixtralProcessor",
1011
+ "tokenizer_class": "TokenizersBackend",
1012
+ "unk_token": "<unk>"
1013
+ }