Claví accent restoration models
Claví restores the accents left out of typed text. Try the models behind it here: one character-level BiLSTM for each of 13 languages, with a standalone Python script that runs on a CPU. Claví is an ai4irl. product.
More about Claví: https://ai4irl.com/clavi
Languages and scores
| Language | Folder | Model version | BiLSTM layers | Hidden size | Embedding size | Vocabulary | Accent-character accuracy | Accent-word accuracy |
|---|---|---|---|---|---|---|---|---|
| Spanish | es-ES |
v3 | 2 | 512 | 128 | 91 | 99.68% | 99.13% |
| Portuguese | pt-PT |
v1 | 2 | 512 | 128 | 103 | 99.63% | 99.05% |
| French | fr-FR |
v6 | 2 | 512 | 128 | 103 | 99.70% | 99.30% |
| German | de-DE |
v2 | 2 | 512 | 128 | 84 | 99.69% | 99.48% |
| Vietnamese | vi-VN |
v3 | 3 | 512 | 192 | 211 | 97.33% | 96.09% |
| Polish | pl-PL |
v1 | 2 | 512 | 128 | 95 | 99.03% | 97.01% |
| Turkish | tr-TR |
v1 | 2 | 512 | 128 | 91 | 99.13% | 97.98% |
| Czech | cs-CZ |
v2 | 2 | 512 | 128 | 107 | 99.40% | 97.81% |
| Romanian | ro-RO |
v2 | 2 | 512 | 128 | 87 | 99.26% | 98.39% |
| Hungarian | hu-HU |
v1 | 2 | 512 | 128 | 95 | 99.35% | 98.49% |
| Italian | it-IT |
v1 | 2 | 512 | 128 | 95 | 99.86% | 99.64% |
| Slovak | sk-SK |
v1 | 2 | 512 | 128 | 111 | 99.45% | 97.85% |
| Dutch | nl-NL |
v1 | 2 | 512 | 128 | 103 | 99.85% | 99.75% |
These scores come from restore_accents.py running on a CPU with the files in this
repository. Each model was evaluated on its recorded validation split, with all
accents removed from the input.
- Accent-character accuracy: the share of letters that can take an accent restored exactly, including letters that had no accent in the original text.
- Accent-word accuracy: the share of words containing such a letter restored exactly.
These splits are not verified to be disjoint from the training text: depending on the language, 8 to 86 validation lines (at most 0.40% of a split) also appear verbatim in the training text.
Quick start
pip install torch safetensors huggingface_hub
hf download ai4irl/clavi-accent-restoration restore_accents.py fr-FR/model.safetensors fr-FR/config.json fr-FR/vocab.json --local-dir clavi
python clavi/restore_accents.py --model-dir clavi/fr-FR
With no text supplied, the script shows the language's demo phrase without accents, then its restoration:
input: Le cafe est deja pret.
restored: Le café est déjà prêt.
To try another language, replace fr-FR in both commands with its
folder name from the table. To restore your own text, add one of --text "...",
--input-file <file> (UTF-8) or --stdin to the Python command.
You can also call the model from Python. Run this inside the clavi folder:
from restore_accents import AccentRestorer
restorer = AccentRestorer("fr-FR")
print(restorer.restore("Le cafe est deja pret."))
The script itself needs only torch and safetensors. huggingface_hub is used
only to download the files.
Demo phrases
Each language folder includes the same demo phrase shown in the Claví app.
Running restore_accents.py --model-dir <language> without text produces the input
and output below. Every phrase must be restored exactly before a release can be
published. These phrases show the models in action; they are not an accuracy test.
For measured accuracy, use the validation scores above.
| Language | Folder | Input (accents removed) | Output |
|---|---|---|---|
| Spanish | es-ES |
Una cancion recien sonada. | Una canción recién soñada. |
| Portuguese | pt-PT |
O avo ficou ao lado da avo. | O avô ficou ao lado da avó. |
| French | fr-FR |
Le cafe est deja pret. | Le café est déjà prêt. |
| German | de-DE |
Schone Gruse aus Munchen. | Schöne Grüße aus München. |
| Vietnamese | vi-VN |
Tieng Viet rat dep. | Tiếng Việt rất đẹp. |
| Polish | pl-PL |
Zazolc gesla jazn. | Zażółć gęślą jaźń. |
| Turkish | tr-TR |
Gunes bugun cok guzel. | Güneş bugün çok güzel. |
| Czech | cs-CZ |
Prilis zlutoucky kun. | Příliš žluťoučký kůň. |
| Romanian | ro-RO |
Cainele asteapta la poarta. | Câinele așteaptă la poartă. |
| Hungarian | hu-HU |
Arvizturo tukorfuro. | Árvíztűrő tükörfúró. |
| Italian | it-IT |
Il nonno e gia arrivato in citta. | Il nonno è già arrivato in città. |
| Slovak | sk-SK |
Dakujem za krasny den. | Ďakujem za krásny deň. |
| Dutch | nl-NL |
Een knusse scene in het cafe. | Een knusse scène in het café. |
Files
restore_accents.py the standalone script above
<language>/model.safetensors the weights, float32
<language>/config.json the architecture (layers, hidden size, embedding size, vocabulary size) and the demo phrase
<language>/vocab.json the characters the model reads and writes, and each letter's accent family
Architecture
- One token per character. Each model's
vocab.jsondefines its vocabulary. Unknown characters are read as<unk>and copied unchanged to the output. - Context in both directions. A character embedding feeds a bidirectional LSTM, followed by a linear layer that scores every vocabulary character at each position. The model reads a whole line in both directions, so surrounding words in that line can inform a word's accents. The table above lists each model's layer count and sizes.
- Same length, same order. Each input character has exactly one output character: nothing is inserted, deleted or reordered. Lines are restored independently.
- Constrained decoding. The highest-scoring output is chosen only from the
input letter's accent family. For example, French
ecan become onlye,é,è,êorë. Letters without an accent family, digits, spaces and punctuation stay unchanged.
Limitations
- Vietnamese (
vi-VN) is the weakest model. Its scores are the lowest in the table. Many syllables combine a vowel mark with a tone mark, and an unaccented spelling can correspond to several different words. - The model is not the app. Claví adds per-language confidence thresholds, waits for a completed word at a word boundary while typing, and writes the correction back into the text field. This script has none of those safeguards or write-back behavior: it chooses an output for every letter regardless of confidence, so it can make corrections the app would hold back.
- The scores have a limited scope. They describe validation text of the same kind the models learned from, with all accents removed. They do not measure performance on names, slang, code, mixed-language text or partly accented text.
- Choose the matching language. Text in another language gets meaningless accents.
Training data
The models were trained on text from FineWeb2 (HuggingFaceFW/fineweb-2),
Wikipedia (wikimedia/wikipedia) and Wikisource (wikimedia/wikisource).
License
Apache License 2.0