Old Church Slavonic Tokenizer and Lemmatizer

This repository provides retrained Stanza-compatible tokenizer and lemmatizer model weights for Old Church Slavonic.

Files

models/
β”œβ”€β”€ new-data/
β”‚   β”œβ”€β”€ tokenize/cu_proiel_tokenizer.pt
β”‚   └── lemma/cu_proiel_nocharlm_lemmatizer.pt
└── combined/
    β”œβ”€β”€ tokenize/cu_proiel_tokenizer.pt
    └── lemma/cu_proiel_nocharlm_lemmatizer.pt

Model variants

New-data model

models/new-data/ contains the tokenizer and lemmatizer trained on the newly annotated Old Church Slavonic dataset.

Combined model

models/combined/ contains the tokenizer and lemmatizer trained on a combined corpus including the newly annotated Old Church Slavonic dataset and UD Old Church Slavonic PROIEL.

License

The model weights are released under CC BY-NC-SA 4.0.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using usmannawaz/old-church-slavonic-tokenizer-lemmatizer 1