urchade commited on
Commit
eac5ac1
·
verified ·
1 Parent(s): 8bc6f48

Upload GLiNER2.5-Decide-1B at step 285000

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ GLiNER-2.5-Decision-HF-Banner.png filter=lfs diff=lfs merge=lfs -text
GLiNER-2.5-Decision-HF-Banner.png ADDED

Git LFS Details

  • SHA256: 4997b9d9e7e84226dcda296409285eb6788567501ab4bec6c6a1d6bff4e1de5d
  • Pointer size: 131 Bytes
  • Size of remote file: 648 kB
README.md ADDED
@@ -0,0 +1,498 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: gliner2
3
+ license: apache-2.0
4
+ language:
5
+ - en
6
+ pipeline_tag: token-classification
7
+ tags:
8
+ - gliner2
9
+ - Text classification
10
+ - Intent classification
11
+ - Sentiment Analysis
12
+ - Topic classification
13
+ - Named Entity Recognition
14
+ base_model: fastino/gliner2-xl-0111
15
+ ---
16
+ <div align="center">
17
+ <a href="https://fastino.ai/lp/gliner?utm_source=huggingface" target="_blank" rel="noopener noreferrer">
18
+ <img src="GLiNER-2.5-Decision-HF-Banner.png" alt="Fastino Agent - Fine-tune GLiNER with a single prompt" width="100%"/>
19
+ </a>
20
+ </div>
21
+
22
+ <div style="display: flex; flex-wrap: wrap; gap: 8px; margin-bottom: 16px;">
23
+ <a href="https://fastino.ai?utm_source=huggingface" target="_blank" rel="noreferrer" style="text-decoration:none;">
24
+ <img src="https://img.shields.io/badge/Finetune-GLiNER2.5-EA4335" alt="Fine-tune and Deploy GLiNER2.5 with Fastino Agent" style="vertical-align:middle;">
25
+ </a>
26
+ <a href="https://arxiv.org/abs/2507.18546" target="_blank" rel="noreferrer" style="text-decoration:none;">
27
+ <img src="https://img.shields.io/badge/arXiv-2507.18546-b31b1b.svg?logo=arxiv" alt="arXiv Paper" style="vertical-align:middle;">
28
+ </a>
29
+ <a href="https://github.com/fastino-ai/GLiNER2" target="_blank" rel="noreferrer" style="text-decoration:none;">
30
+ <img src="https://img.shields.io/badge/GitHub-GLiNER2-black?logo=github" alt="GitHub" style="vertical-align:middle;">
31
+ </a>
32
+ <a href="https://x.com/fastinoAI" target="_blank" rel="noreferrer" style="text-decoration:none;">
33
+ <img src="https://img.shields.io/twitter/follow/:fastinoAI" alt="Follow @fastinoAI" style="vertical-align:middle;">
34
+ </a>
35
+ </div>
36
+
37
+ # GLiNER2.5-Decide-1B
38
+
39
+ **The 1B classification model in the GLiNER2.5 family.** Pass any label set at call time: intent, routing, sentiment, priority, policy, and multi-label tags, in a single forward pass. No prompt template. No generated tokens. Load it with `AutoExtractor` and ship it locally.
40
+
41
+ An earlier 1B checkpoint at step 195000 scored 59.6% on the 17-domain held-out suite, ahead of open typed-decision baselines up to 4B. This release uploads the newer **step-285000** weights; that exact checkpoint has not yet been re-evaluated. A single call can score several heads at once. Single-label tasks return one string. Multi-label tasks return every label above the threshold.
42
+
43
+ This release is not a general-purpose model. It does not reason, explain, or answer open questions. It was not trained on public benchmarks. It is a specialist for operational decisions: customer and banking intent, travel and clinic requests, review sentiment, document type, email and ticket routing, human handoff, agent completion, moderation, severity, urgency, and spam. The same call also answers a question about a passage, classifies a book, uses labels that carry a description, and scores an ordinal scale.
44
+
45
+ Outputs below are potential results for these inputs. They show the shape `classify_text` returns.
46
+
47
+ ## Install
48
+
49
+ ```bash
50
+ pip install gliner2
51
+ ```
52
+
53
+ ```python
54
+ from gliner2 import AutoExtractor
55
+
56
+ model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide-1B")
57
+ ```
58
+
59
+ ## Examples
60
+
61
+ ### Customer support intent
62
+
63
+ Route a live message before a human ever sees it. Refunds, cancellations, login failures, and shipping delays share one inbox and one label set. Decide picks the action the queue should take, so the right workflow starts on the first turn.
64
+
65
+ ```python
66
+ model.classify_text(
67
+ "My subscription renewed on April 15 for ¥5,400 after the service was already down. Can I get that charge refunded?",
68
+ {"intent": [
69
+ "order_status", "refund_request", "cancel_subscription", "update_payment",
70
+ "login_problem", "shipping_delay", "bug_report", "speak_to_human", "other",
71
+ ]},
72
+ )
73
+ ```
74
+
75
+ Potential output:
76
+
77
+ ```text
78
+ {"intent": "refund_request"}
79
+ ```
80
+
81
+ ### Banking request
82
+
83
+ A single customer message often mixes a failed payment, a beneficiary change, and a fee question. The label set is the product catalog: transfers, cards, fraud, mortgages. The model maps the utterance onto the operation the core system should open.
84
+
85
+ ```python
86
+ model.classify_text(
87
+ "The transfer I sent this morning is still pending, and I think I used the wrong sort code. Can you stop it and add Emily as the beneficiary instead?",
88
+ {"intent": [
89
+ "transfer_pending", "transfer_cancel", "beneficiary_add", "card_lost",
90
+ "balance_inquiry", "fraud_report", "mortgage_application", "fee_explanation",
91
+ ]},
92
+ )
93
+ ```
94
+
95
+ Potential output:
96
+
97
+ ```text
98
+ {"intent": "transfer_cancel"}
99
+ ```
100
+
101
+ ### Travel request
102
+
103
+ Booking, changing, cancelling, and seat requests look similar in free text and trigger different inventory calls. Use this when a chat or email has to become a structured booking action without a form.
104
+
105
+ ```python
106
+ model.classify_text(
107
+ "I need to move my Friday flight to Paris to Saturday morning, same cabin, and keep the aisle seat if you can.",
108
+ {"request": ["book", "change", "cancel", "status", "seat_change", "refund", "baggage"]},
109
+ )
110
+ ```
111
+
112
+ Potential output:
113
+
114
+ ```text
115
+ {"request": "change"}
116
+ ```
117
+
118
+ ### Clinic request
119
+
120
+ Patients describe symptoms and the thing they want in the same sentence: an appointment, a refill, a result, a referral. Front-desk routing needs that distinction before anyone is put on a schedule.
121
+
122
+ ```python
123
+ model.classify_text(
124
+ "The rash came back after the antibiotics finished. Can I get a same-week appointment with dermatology, or should I just refill the cream?",
125
+ {"request": [
126
+ "book_appointment", "refill_prescription", "test_results",
127
+ "referral", "billing_question", "cancel_appointment",
128
+ ]},
129
+ )
130
+ ```
131
+
132
+ Potential output:
133
+
134
+ ```text
135
+ {"request": "book_appointment"}
136
+ ```
137
+
138
+ ### Review sentiment
139
+
140
+ Star ratings hide mixed reviews. A product can be praised and rejected in one paragraph. A four-way sentiment label is what a dashboard, a reply policy, or a ranking feature actually needs.
141
+
142
+ ```python
143
+ model.classify_text(
144
+ "Battery dies before lunch, but the keyboard and the screen are the best I have used on a laptop.",
145
+ {"sentiment": ["positive", "negative", "mixed", "neutral"]},
146
+ )
147
+ ```
148
+
149
+ Potential output:
150
+
151
+ ```text
152
+ {"sentiment": "mixed"}
153
+ ```
154
+
155
+ ### Product aspects
156
+
157
+ Sentiment says the review is mixed. Aspects say why: battery, keyboard, screen. Several labels apply at once, so this head is multi-label. That is the input to aspect-level analytics and to a reply that mentions the right part of the product.
158
+
159
+ ```python
160
+ model.classify_text(
161
+ "Battery dies before lunch, but the keyboard and the screen are the best I have used on a laptop.",
162
+ {"aspects": {
163
+ "labels": ["battery", "keyboard", "screen", "camera", "price", "support"],
164
+ "multi_label": True,
165
+ "cls_threshold": 0.4,
166
+ }},
167
+ )
168
+ ```
169
+
170
+ Potential output:
171
+
172
+ ```text
173
+ {"aspects": ["battery", "keyboard", "screen"]}
174
+ ```
175
+
176
+ ### News topic
177
+
178
+ Wire copy, alerts, and scraped headlines have to land in a section before they are ranked or summarized. The label set is the section list of the product, not a fixed taxonomy baked into the model.
179
+
180
+ ```python
181
+ model.classify_text(
182
+ "The central bank held rates and said inflation is still above target, pushing bank stocks lower in afternoon trading.",
183
+ {"topic": ["politics", "business", "sports", "science", "entertainment", "world"]},
184
+ )
185
+ ```
186
+
187
+ Potential output:
188
+
189
+ ```text
190
+ {"topic": "business"}
191
+ ```
192
+
193
+ ### Document type
194
+
195
+ Inboxes and shared drives mix invoices, contracts, resumes, and notes. Classifying the document is the gate in front of extraction: an invoice goes to payables, a contract goes to review, a resume goes to screening.
196
+
197
+ ```python
198
+ model.classify_text(
199
+ "INVOICE 1842\nBill to: Northstar QA\nAmount due: 2,400 USD\nDue: 30 April 2026\nWire instructions are on page 2.",
200
+ {"document_type": ["invoice", "receipt", "contract", "resume", "support_email", "meeting_notes"]},
201
+ )
202
+ ```
203
+
204
+ Potential output:
205
+
206
+ ```text
207
+ {"document_type": "invoice"}
208
+ ```
209
+
210
+ ### Email triage
211
+
212
+ A shared mailbox needs three answers before a message is filed: what the sender wants, how soon it matters, and which team owns it. One call scores all three heads on the same text, so the router does not run the model three times.
213
+
214
+ ```python
215
+ model.classify_text(
216
+ "From: compliance@group.example\nSubject: Protocol update — action required today\n\nPlease confirm the new retention rule is applied before Friday's audit.",
217
+ {
218
+ "intent": ["fyi", "request", "approval", "complaint", "newsletter", "security_alert"],
219
+ "urgency": ["low", "normal", "high", "critical"],
220
+ "route": ["support", "billing", "legal", "security", "finance", "archive"],
221
+ },
222
+ )
223
+ ```
224
+
225
+ Potential output:
226
+
227
+ ```text
228
+ {"intent": "request", "urgency": "high", "route": "legal"}
229
+ ```
230
+
231
+ ### Ticket routing
232
+
233
+ Employees describe a problem, not a department. Payroll, benefits, IT access, and facilities share the same portal. The queue label is the assignment, so the ticket opens in the right team instead of bouncing through a dispatcher.
234
+
235
+ ```python
236
+ model.classify_text(
237
+ "[subject] 401k deduction missing from this paystub\n[body] Last month's contribution posted. This month the line is gone and HR told me to open a ticket.",
238
+ {"queue": [
239
+ "payroll", "benefits", "it_access", "facilities",
240
+ "expense_reimbursement", "manager_approval",
241
+ ]},
242
+ )
243
+ ```
244
+
245
+ Potential output:
246
+
247
+ ```text
248
+ {"queue": "benefits"}
249
+ ```
250
+
251
+ ### Handoff to a person
252
+
253
+ Most turns should stay automated. A repeated complaint, an explicit request for a human, or a case the bot cannot close should leave the flow. This is the yes/no gate in front of an agent queue.
254
+
255
+ ```python
256
+ model.classify_text(
257
+ "This is the third time I have explained the same missing refund. Stop the bot and get me a person.",
258
+ {"handoff": ["yes", "no"]},
259
+ )
260
+ ```
261
+
262
+ Potential output:
263
+
264
+ ```text
265
+ {"handoff": "yes"}
266
+ ```
267
+
268
+ ### Did the agent finish?
269
+
270
+ A trace can look busy and still be incomplete: a draft saved, a button disabled, a required field empty. Supervisors and eval harnesses need a finish decision from the goal plus the last state, not from whether the model stopped talking.
271
+
272
+ ```python
273
+ model.classify_text(
274
+ "Goal: email the Q4 summary to every partner.\nLast action: draft saved in the hub.\nSend button is still disabled because two partners have no address.",
275
+ {"finished": ["yes", "no"]},
276
+ )
277
+ ```
278
+
279
+ Potential output:
280
+
281
+ ```text
282
+ {"finished": "no"}
283
+ ```
284
+
285
+ ### Moderation
286
+
287
+ User content and model output both need a policy decision before they are shown or acted on. The labels are the policy, from allow through personal data, harassment, scam, and spam, so a filter can block, redact, or escalate.
288
+
289
+ ```python
290
+ model.classify_text(
291
+ "Post the customer's home address in the public thread so everyone can see where the package actually went.",
292
+ {"policy": ["allow", "personal_data", "harassment", "scam", "violence", "spam"]},
293
+ )
294
+ ```
295
+
296
+ Potential output:
297
+
298
+ ```text
299
+ {"policy": "personal_data"}
300
+ ```
301
+
302
+ ### Incident severity
303
+
304
+ Pages and deploys produce more notes than pages. Severity is what decides whether this wakes someone up. Staging-only tag drift is not a production checkout failure, and the label should say so.
305
+
306
+ ```python
307
+ model.classify_text(
308
+ "The deploy left resource tags inconsistent across staging. Production checkout is unaffected. No customer reports yet.",
309
+ {"severity": ["info", "low", "medium", "high", "critical"]},
310
+ )
311
+ ```
312
+
313
+ Potential output:
314
+
315
+ ```text
316
+ {"severity": "low"}
317
+ ```
318
+
319
+ ### Urgency score
320
+
321
+ Some queues want a rank, not a bucket. Pass `"0"` through `"5"` as ordinary strings. A payroll cutoff before 5pm and a typo in a wiki page should not receive the same score, and the label is what the SLA system reads.
322
+
323
+ ```python
324
+ model.classify_text(
325
+ "Payroll file has to be corrected before the 5pm cutoff or the whole company is paid late.",
326
+ {"urgency": ["0", "1", "2", "3", "4", "5"]},
327
+ )
328
+ ```
329
+
330
+ Potential output:
331
+
332
+ ```text
333
+ {"urgency": "5"}
334
+ ```
335
+
336
+ ### Spam or not
337
+
338
+ The first filter on an inbox or a comment stream. Phishing and mailbox-full lures should never reach the intent router. A two-label decision keeps that check cheap enough to run on every message.
339
+
340
+ ```python
341
+ model.classify_text(
342
+ "Your mailbox is almost full. Click here in the next hour or we will delete every message.",
343
+ {"label": ["spam", "ham"]},
344
+ )
345
+ ```
346
+
347
+ Potential output:
348
+
349
+ ```text
350
+ {"label": "spam"}
351
+ ```
352
+
353
+ ### Several decisions at once
354
+
355
+ Real requests carry more than one fact. A guest can need a room move, a billing release, and a maintenance ticket in the same call. Intent, priority, a human gate, and multi-label topics are scored together, which is how a property system opens the right work orders without a chain of prompts.
356
+
357
+ ```python
358
+ model.classify_text(
359
+ "Guest in room 1408 says the AC has been out since yesterday and they want to move tonight or leave. They also asked for the incidentals hold to be released.",
360
+ {
361
+ "intent": ["maintenance", "room_change", "checkout", "billing", "complaint", "amenity_request"],
362
+ "priority": ["low", "normal", "high", "urgent"],
363
+ "needs_human": ["yes", "no"],
364
+ "topics": {
365
+ "labels": ["hvac", "billing", "housekeeping", "noise", "safety"],
366
+ "multi_label": True,
367
+ "cls_threshold": 0.4,
368
+ },
369
+ },
370
+ )
371
+ ```
372
+
373
+ Potential output:
374
+
375
+ ```text
376
+ {
377
+ "intent": "room_change",
378
+ "priority": "high",
379
+ "needs_human": "yes",
380
+ "topics": ["hvac", "billing"]
381
+ }
382
+ ```
383
+
384
+ ### Question over a passage
385
+
386
+ The text is the source. The question is the task. Yes or no is the whole answer: no chain of thought, no extracted sentence, just the decision a checker or a form needs.
387
+
388
+ ```python
389
+ model.classify_text(
390
+ "The treaty was signed in Paris in 1992. It entered into force the following year, after the last signatory ratified it.",
391
+ {"answer": {
392
+ "labels": ["yes", "no"],
393
+ "prompt": "Did the treaty enter into force in 1992?",
394
+ }},
395
+ )
396
+ ```
397
+
398
+ Potential output:
399
+
400
+ ```text
401
+ {"answer": "no"}
402
+ ```
403
+
404
+ ### Book
405
+
406
+ A catalog, a slush pile, or a library inbox needs a genre before anyone writes a blurb. The passage is enough. The label set is the shelf list.
407
+
408
+ ```python
409
+ model.classify_text(
410
+ "She closed the ledger, blew out the lamp, and listened for the stair. The house had been empty since the winter the river took the bridge.",
411
+ {"genre": ["mystery", "romance", "history", "science_fiction", "literary_fiction", "cookbook"]},
412
+ )
413
+ ```
414
+
415
+ Potential output:
416
+
417
+ ```text
418
+ {"genre": "literary_fiction"}
419
+ ```
420
+
421
+ ### Labels with a description
422
+
423
+ When the name of a label is not enough, pass a short description with it. The description is part of the decision, which is how a private taxonomy stays precise without a bigger model.
424
+
425
+ ```python
426
+ model.classify_text(
427
+ "Please reset the card PIN. The new one never arrived and the old one is locked after three tries.",
428
+ {"intent": {
429
+ "labels": {
430
+ "card_pin_change": "The customer wants a new PIN or the current PIN replaced",
431
+ "card_lost": "The physical card is missing",
432
+ "balance_inquiry": "The customer wants the current balance",
433
+ },
434
+ }},
435
+ )
436
+ ```
437
+
438
+ Potential output:
439
+
440
+ ```text
441
+ {"intent": "card_pin_change"}
442
+ ```
443
+
444
+ ### Ordinal score
445
+
446
+ Some products want a rank, not a class. Pass the scale as ordinary strings, from `"0"` to `"10"`. A review, a rubric, or a satisfaction form can read the label directly.
447
+
448
+ ```python
449
+ model.classify_text(
450
+ "I finished it in two nights. The ending is earned, the middle drags, and I would still hand it to a friend.",
451
+ {"rating": ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10"]},
452
+ )
453
+ ```
454
+
455
+ Potential output:
456
+
457
+ ```text
458
+ {"rating": "7"}
459
+ ```
460
+
461
+ ## Benchmark
462
+
463
+ Exact-match accuracy on [`fastino/fast-decisions`](https://huggingface.co/datasets/fastino/fast-decisions): 17 domains, 300 held-out examples each, the same text and the same candidate labels for every model.
464
+
465
+ | Model | Avg |
466
+ |---|---:|
467
+ | GLiNER2.5-Decide (340M) | **60.2%** |
468
+ | **GLiNER2.5-Decide-1B (step 195000)** | **59.6%** |
469
+ | JevK5 | 57.6% |
470
+ | SemIf (Qwen3.5-4B) | 56.4% |
471
+ | GLiFormer large-v1 | 49.0% |
472
+ | Laya Router | 46.6% |
473
+
474
+ The reported 1B result is from step 195000; this repository contains newer step-285000 weights. Domain list and label sets are on the [dataset card](https://huggingface.co/datasets/fastino/fast-decisions).
475
+
476
+ ## Details
477
+
478
+ - **What it is:** a specialist classifier for operational decisions, including questions over a passage, books, described labels, and ordinal scores
479
+ - **What it is not:** a general-purpose model. No reasoning, no explanations, no open-ended answers. Not trained on public benchmarks.
480
+ - **Encoder:** Ettin encoder-from-decoder 1B (`jhu-clsp/ettin-enc-from-dec-1b`)
481
+ - **Parameters:** approximately 1B
482
+ - **Checkpoint:** step 285000
483
+ - **Runs on:** CPU or GPU, through `gliner2`
484
+ - **License:** Apache 2.0
485
+
486
+ ## Citation
487
+
488
+ ```bibtex
489
+ @misc{zaratiana2025gliner2efficientmultitaskinformation,
490
+ title={GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface},
491
+ author={Urchade Zaratiana and Gil Pasternak and Oliver Boyd and George Hurn-Maloney and Ash Lewis},
492
+ year={2025},
493
+ eprint={2507.18546},
494
+ archivePrefix={arXiv},
495
+ primaryClass={cs.CL},
496
+ url={https://arxiv.org/abs/2507.18546},
497
+ }
498
+ ```
config.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architecture": "span",
3
+ "architecture_version": 1,
4
+ "architectures": [
5
+ "SpanExtractor"
6
+ ],
7
+ "attn_implementation": "sdpa",
8
+ "config_version": 3,
9
+ "counting_layer": "count_lstm",
10
+ "max_len": null,
11
+ "max_width": 8,
12
+ "model_name": "jhu-clsp/ettin-enc-from-dec-1b",
13
+ "model_type": "extractor",
14
+ "span_head": {
15
+ "dropout": 0.1,
16
+ "max_width": 8,
17
+ "span_mode": "markerV0"
18
+ },
19
+ "token_pooling": "first",
20
+ "transformers_version": "5.17.0"
21
+ }
encoder_config/config.json ADDED
@@ -0,0 +1,87 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "ModernBertForMaskedLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 50281,
8
+ "causal_mask": false,
9
+ "classifier_activation": "gelu",
10
+ "classifier_bias": false,
11
+ "classifier_dropout": 0.0,
12
+ "classifier_pooling": "mean",
13
+ "cls_token_id": 50281,
14
+ "decoder_bias": true,
15
+ "deterministic_flash_attn": false,
16
+ "dtype": "float32",
17
+ "embedding_dropout": 0.0,
18
+ "eos_token_id": 50282,
19
+ "global_attn_every_n_layers": 3,
20
+ "gradient_checkpointing": false,
21
+ "hidden_activation": "gelu",
22
+ "hidden_size": 1792,
23
+ "initializer_cutoff_factor": 2.0,
24
+ "initializer_range": 0.02,
25
+ "intermediate_size": 3840,
26
+ "is_causal": false,
27
+ "layer_norm_eps": 1e-05,
28
+ "layer_types": [
29
+ "full_attention",
30
+ "sliding_attention",
31
+ "sliding_attention",
32
+ "full_attention",
33
+ "sliding_attention",
34
+ "sliding_attention",
35
+ "full_attention",
36
+ "sliding_attention",
37
+ "sliding_attention",
38
+ "full_attention",
39
+ "sliding_attention",
40
+ "sliding_attention",
41
+ "full_attention",
42
+ "sliding_attention",
43
+ "sliding_attention",
44
+ "full_attention",
45
+ "sliding_attention",
46
+ "sliding_attention",
47
+ "full_attention",
48
+ "sliding_attention",
49
+ "sliding_attention",
50
+ "full_attention",
51
+ "sliding_attention",
52
+ "sliding_attention",
53
+ "full_attention",
54
+ "sliding_attention",
55
+ "sliding_attention",
56
+ "full_attention"
57
+ ],
58
+ "local_attention": 128,
59
+ "masked_prediction": true,
60
+ "max_position_embeddings": 7999,
61
+ "mlp_bias": false,
62
+ "mlp_dropout": 0.0,
63
+ "model_type": "modernbert",
64
+ "norm_bias": false,
65
+ "norm_eps": 1e-05,
66
+ "num_attention_heads": 28,
67
+ "num_hidden_layers": 28,
68
+ "pad_token_id": 50283,
69
+ "position_embedding_type": "sans_pos",
70
+ "repad_logits_with_grad": false,
71
+ "rope_parameters": {
72
+ "full_attention": {
73
+ "rope_theta": 160000.0,
74
+ "rope_type": "default"
75
+ },
76
+ "sliding_attention": {
77
+ "rope_theta": 160000.0,
78
+ "rope_type": "default"
79
+ }
80
+ },
81
+ "sep_token_id": 50282,
82
+ "sparse_pred_ignore_index": -100,
83
+ "sparse_prediction": false,
84
+ "tie_word_embeddings": true,
85
+ "transformers_version": "5.17.0",
86
+ "vocab_size": 50378
87
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:02c567d791aed26550d300064c7f0c0094fd65291503c65969b45b30786e33b3
3
+ size 4755208228
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "clean_up_tokenization_spaces": true,
4
+ "cls_token": "[CLS]",
5
+ "extra_special_tokens": [
6
+ "[SEP_STRUCT]",
7
+ "[SEP_TEXT]",
8
+ "[P]",
9
+ "[C]",
10
+ "[E]",
11
+ "[R]",
12
+ "[L]",
13
+ "[EXAMPLE]",
14
+ "[OUTPUT]",
15
+ "[DESCRIPTION]"
16
+ ],
17
+ "is_local": false,
18
+ "local_files_only": false,
19
+ "mask_token": "[MASK]",
20
+ "model_input_names": [
21
+ "input_ids",
22
+ "attention_mask"
23
+ ],
24
+ "model_max_length": 8192,
25
+ "pad_token": "[PAD]",
26
+ "sep_token": "[SEP]",
27
+ "tokenizer_class": "TokenizersBackend",
28
+ "unk_token": "[UNK]"
29
+ }