--- license: apache-2.0 language: - ne - en tags: - tokenizer - nepali - devanagari - unicode - trie - nlp - tokenization --- # Supernova NepaliFast V4 Supernova NepaliFast V4 is a Nepali-first, Unicode-aware Longest-Match Trie tokenizer. ## Features | Feature | Status | |---|---| | Nepali-first | PASS | | Devanagari | PASS | | English | PASS | | Unicode | PASS | | Emoji | PASS | | Mathematical symbols | PASS | | Multilingual text | PASS | | Round-trip decoding | PASS | | CPU-friendly | PASS | ## Vocabulary - Vocabulary size: 2,890 - ID range: 0 -> 2889 - ID integrity: PASS ## Final Extreme Benchmark - Documents: 9,120 - Characters: 8,597,880 - Unknown characters: 0 - Round-trip failures: 0 - Fallback documents: 0 | Engine | Characters/sec | Tokens/sec | |---|---:|---:| | Supernova V4 | 7,886,249 | 6,897,069 | | Tiktoken o200k | 7,666,757 | 4,802,143 | ### Relative performance - Character throughput: 1.03x - Token throughput: 1.44x ## Tested Unicode ```text √2 ≈ 1.4142135623730951 ∑(xᵢ²) → ∞ 🇳🇵 🚀 🔥 🤖 🧠 💻 🌋 👨‍👩‍👧‍👦 👩‍💻 🧑‍🚀 — – … « » “ ” ‘ ’ ≠ ≤ ≥ ± × ÷ ∞ ``` ## Nepali ```text नमस्ते नेपाल लुम्बिनी नेपालको प्रसिद्ध स्थान हो। सगरमाथा नेपालको गौरव हो। लाख करोड अरब खर्ब हजार ``` ## Run on your own computer Install Python 3.9 or newer. Run the included benchmark: ```bash python benchmark.py ``` The repository contains the tokenizer vocabulary and a reference Python implementation for testing. ## Research Focus - Nepali-first tokenization - Devanagari coverage - Unicode robustness - Deterministic tokenization - Lossless round-trip decoding - High token throughput - CPU-friendly execution Supernova NepaliFast V4 is a tokenizer, not a language model. ## License Apache License 2.0. ## Supernova AI Built as part of the Supernova AI tokenizer research project. **Fast. Unicode-safe. Nepali-first.**