Bulgarian small-model toolkit
Small models doing narrow tasks, trained from scratch: a 91M Bulgarian LM, a shlyokavitsa restorer (model + live demo), and the dataset behind it.
Text Generation • 91.3M • Updated • 52Note The 91.26M flagship - ties Gemini 3.5-flash at judging Bulgarian fluency, loses on raw compression.
glassbox/gpt-alpha-bg-restorer
Text Generation • 4.73M • Updated • 54Note Two character-level GPTs (3.16M / 4.73M) restoring shlyokavitsa to Cyrillic - beats a 2.6B general model at the task by a category, not a margin.
glassbox/shlyokavitsa-pairs
Viewer • Updated • 210k • 83Note 210k (Latin, Cyrillic) pairs from Bulgarian Wikipedia - the first shlyokavitsa dataset on the Hub.
Bulgarian Shlyokavitsa Restorer
⌨Latin-typed Bulgarian back into Cyrillic, in-browser
Note The restorer, live in your browser - type shlyokavitsa, watch it become Cyrillic, 100% client-side.
glassbox/gpt-alpha-bg-14m-onnx
Text Generation • 13.8M • Updated • 150Note The 13.77M LM as 17 MB of int8 ONNX, running client-side in the site's /live-lm page; the quantization cost is measured (+0.0018 bpc) and the fp32 graph is verified against PyTorch at 1.7e-5 max logit difference.