--- license: mit pipeline_tag: text-classification library_name: keras tags: - language-identification language: - ru - es - fr - de --- # langmonster-1.0 Sentence-level language identification for 19 languages, trained on Tatoeba with a compact transformer encoder. ## Model details | | | |---|---| | Architecture | Transformer encoder (3 blocks, d_model 256, 8 heads, FFN 1024) + global average pooling + softmax | | Parameters | 15.17M | | Tokenizer | Word-level `TextVectorization`, vocab 50k, fixed `maxlen=128` (no char n-grams — static shapes for TPU) | | Input | Sentence as token ids, int32, shape `(128,)` | | Output | Softmax logits over 19 languages | | Framework | TensorFlow 2.20 / Keras 3 | ## Supported languages eng, spa, fra, deu, rus, bul, ell, ara, cmn, jpn, ita, por, tur, nld, hin, pol, tha, urd, vie ## Training data Tatoeba sentence export (`sentences.csv`, 13.5M sentences), filtered to the 19 languages above and capped per language. ``` valid sentences: 2,870,490 (19 languages) eng: 250,000 ita: 250,000 spa: 250,000 por: 250,000 fra: 250,000 tur: 250,000 deu: 250,000 nld: 200,710 rus: 250,000 pol: 137,076 jpn: 248,866 cmn: 88,787 ara: 68,485 ell: 42,264 vie: 32,430 bul: 25,698 hin: 16,475 tha: 6,848 urd: 2,851 ``` ## Training procedure - Optimizer: Adam, lr 1e-3 - Global batch size 512, 3 epochs (18,575 steps/epoch on the repeated dataset) - Hardware: NVIDIA A100 40GB - Final: train accuracy **0.9692**, loss 0.0915 **Performance expectations:** strong on the 5 well-represented languages (eng/spa/fra/deu/rus, 250k each). Expect worse results on low-resource languages (urd, tha, hin, bul) and on domains far from Tatoeba text. Reported accuracy is on the training split; there is no held-out validation set for this release. ## Usage ```bash python infer.py --model artifacts/model.keras \ "text on english" "русский текст" "今日の天気は?" ``` Expected outputs: `text on english` -> eng (1.000), `русский текст` -> rus (0.967). Programmatic (Keras): ```python import json import numpy as np import tensorflow as tf lang_map = json.load(open("artifacts/languages.json")) index_to_lang = {v: k for k, v in lang_map.items()} vocab = json.load(open("artifacts/vocab.json")) vectorizer = tf.keras.layers.TextVectorization( max_tokens=len(vocab), output_mode="int", output_sequence_length=128, standardize="lower_and_strip_punctuation", ) vectorizer.set_vocabulary(vocab) model = tf.keras.models.load_model("artifacts/model.keras") tokens = tf.cast(vectorizer(tf.constant(["hello world"])), tf.int32) logits = model.predict(tokens) print(index_to_lang[int(np.argmax(logits[0]))]) ``` ## Limitations - Trained only on Tatoeba sentences; informal/social text and other domains may degrade accuracy. - Tatoeba has no Swahili sentences, so swa is excluded despite being in the original target list. - Very low-resource classes (urd, tha, hin, bul) have only thousands of samples and are unreliable. - Sentence-level only; longer documents should be split into sentences before classification. ## License MIT. Data is from [Tatoeba](https://tatoeba.org), whose sentences are subject to their own [licensing terms](https://tatoeba.org/en/terms_of_use).