PyTorch
Safetensors
accent-restoration
diacritics-restoration
character-level
bilstm

Claví accent restoration models

Claví restores the accents left out of typed text. Try the models behind it here: one character-level BiLSTM for each of 13 languages, with a standalone Python script that runs on a CPU. Claví is an ai4irl. product.

More about Claví: https://ai4irl.com/clavi

Languages and scores

Language Folder Model version BiLSTM layers Hidden size Embedding size Vocabulary Accent-character accuracy Accent-word accuracy
Spanish es-ES v3 2 512 128 91 99.68% 99.13%
Portuguese pt-PT v1 2 512 128 103 99.63% 99.05%
French fr-FR v6 2 512 128 103 99.70% 99.30%
German de-DE v2 2 512 128 84 99.69% 99.48%
Vietnamese vi-VN v3 3 512 192 211 97.33% 96.09%
Polish pl-PL v1 2 512 128 95 99.03% 97.01%
Turkish tr-TR v1 2 512 128 91 99.13% 97.98%
Czech cs-CZ v2 2 512 128 107 99.40% 97.81%
Romanian ro-RO v2 2 512 128 87 99.26% 98.39%
Hungarian hu-HU v1 2 512 128 95 99.35% 98.49%
Italian it-IT v1 2 512 128 95 99.86% 99.64%
Slovak sk-SK v1 2 512 128 111 99.45% 97.85%
Dutch nl-NL v1 2 512 128 103 99.85% 99.75%

These scores come from restore_accents.py running on a CPU with the files in this repository. Each model was evaluated on its recorded validation split, with all accents removed from the input.

  • Accent-character accuracy: the share of letters that can take an accent restored exactly, including letters that had no accent in the original text.
  • Accent-word accuracy: the share of words containing such a letter restored exactly.

These splits are not verified to be disjoint from the training text: depending on the language, 8 to 86 validation lines (at most 0.40% of a split) also appear verbatim in the training text.

Quick start

pip install torch safetensors huggingface_hub
hf download ai4irl/clavi-accent-restoration restore_accents.py fr-FR/model.safetensors fr-FR/config.json fr-FR/vocab.json --local-dir clavi
python clavi/restore_accents.py --model-dir clavi/fr-FR

With no text supplied, the script shows the language's demo phrase without accents, then its restoration:

input:    Le cafe est deja pret.
restored: Le café est déjà prêt.

To try another language, replace fr-FR in both commands with its folder name from the table. To restore your own text, add one of --text "...", --input-file <file> (UTF-8) or --stdin to the Python command.

You can also call the model from Python. Run this inside the clavi folder:

from restore_accents import AccentRestorer

restorer = AccentRestorer("fr-FR")
print(restorer.restore("Le cafe est deja pret."))

The script itself needs only torch and safetensors. huggingface_hub is used only to download the files.

Demo phrases

Each language folder includes the same demo phrase shown in the Claví app. Running restore_accents.py --model-dir <language> without text produces the input and output below. Every phrase must be restored exactly before a release can be published. These phrases show the models in action; they are not an accuracy test. For measured accuracy, use the validation scores above.

Language Folder Input (accents removed) Output
Spanish es-ES Una cancion recien sonada. Una canción recién soñada.
Portuguese pt-PT O avo ficou ao lado da avo. O avô ficou ao lado da avó.
French fr-FR Le cafe est deja pret. Le café est déjà prêt.
German de-DE Schone Gruse aus Munchen. Schöne Grüße aus München.
Vietnamese vi-VN Tieng Viet rat dep. Tiếng Việt rất đẹp.
Polish pl-PL Zazolc gesla jazn. Zażółć gęślą jaźń.
Turkish tr-TR Gunes bugun cok guzel. Güneş bugün çok güzel.
Czech cs-CZ Prilis zlutoucky kun. Příliš žluťoučký kůň.
Romanian ro-RO Cainele asteapta la poarta. Câinele așteaptă la poartă.
Hungarian hu-HU Arvizturo tukorfuro. Árvíztűrő tükörfúró.
Italian it-IT Il nonno e gia arrivato in citta. Il nonno è già arrivato in città.
Slovak sk-SK Dakujem za krasny den. Ďakujem za krásny deň.
Dutch nl-NL Een knusse scene in het cafe. Een knusse scène in het café.

Files

restore_accents.py             the standalone script above
<language>/model.safetensors   the weights, float32
<language>/config.json         the architecture (layers, hidden size, embedding size, vocabulary size) and the demo phrase
<language>/vocab.json          the characters the model reads and writes, and each letter's accent family

Architecture

  • One token per character. Each model's vocab.json defines its vocabulary. Unknown characters are read as <unk> and copied unchanged to the output.
  • Context in both directions. A character embedding feeds a bidirectional LSTM, followed by a linear layer that scores every vocabulary character at each position. The model reads a whole line in both directions, so surrounding words in that line can inform a word's accents. The table above lists each model's layer count and sizes.
  • Same length, same order. Each input character has exactly one output character: nothing is inserted, deleted or reordered. Lines are restored independently.
  • Constrained decoding. The highest-scoring output is chosen only from the input letter's accent family. For example, French e can become only e, é, è, ê or ë. Letters without an accent family, digits, spaces and punctuation stay unchanged.

Limitations

  • Vietnamese (vi-VN) is the weakest model. Its scores are the lowest in the table. Many syllables combine a vowel mark with a tone mark, and an unaccented spelling can correspond to several different words.
  • The model is not the app. Claví adds per-language confidence thresholds, waits for a completed word at a word boundary while typing, and writes the correction back into the text field. This script has none of those safeguards or write-back behavior: it chooses an output for every letter regardless of confidence, so it can make corrections the app would hold back.
  • The scores have a limited scope. They describe validation text of the same kind the models learned from, with all accents removed. They do not measure performance on names, slang, code, mixed-language text or partly accented text.
  • Choose the matching language. Text in another language gets meaningless accents.

Training data

The models were trained on text from FineWeb2 (HuggingFaceFW/fineweb-2), Wikipedia (wikimedia/wikipedia) and Wikisource (wikimedia/wikisource).

License

Apache License 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train ai4irl/clavi-accent-restoration