mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
Paper • 2509.06888 • Published • 15
mmBERT is trained on 3T tokens from over 1800 languages, showing SoTA scores on benchmarks and exceptional low-resource performance
Note ICML 2026
Note Intermediate checkpoints for continued pre-training (MosiacML Composer format)
Note Pre-training Data
Note Randomized data for training (not recommended unless you are using the same data mix)