File size: 4,503 Bytes
2930bce
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4e299c6
2930bce
4e299c6
 
2930bce
 
 
 
 
 
 
 
 
 
 
 
4e299c6
2930bce
 
 
 
 
 
4e299c6
 
2930bce
 
 
 
 
4e299c6
 
2930bce
 
 
 
 
 
 
 
 
 
4e299c6
 
 
2930bce
 
 
 
4e299c6
2930bce
 
 
 
 
 
 
 
 
 
 
4e299c6
2930bce
4e299c6
 
2930bce
7e91909
 
 
 
 
 
 
2930bce
 
7e91909
 
 
 
2930bce
7e91909
 
 
 
 
 
 
2930bce
 
7e91909
 
2930bce
 
 
4e299c6
 
2930bce
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
---
license: cc-by-sa-4.0
language:
  - te
tags:
  - dependency-parsing
  - combo
  - universal-dependencies
datasets:
  - universal_dependencies
model-name: Combo Nlp Xlm Roberta Base Telugu Mtg Ud2.17
pipeline_tag: token-classification
---

# COMBO-NLP Model for Telugu

## Model Description

This is a Telugu-language model based on [combo-nlp](https://pypi.org/project/combo-nlp/), an open-source natural language preprocessing system. It performs:

- sentence segmentation (via [combo-seg](https://pypi.org/project/combo-seg/))
- tokenisation (via [combo-seg](https://pypi.org/project/combo-seg/))
- part-of-speech tagging
- morphological analysis
- lemmatisation
- dependency parsing

The Telugu model uses ``FacebookAI/xlm-roberta-base`` as its base encoder and is trained on [UD_Telugu-MTG](https://github.com/UniversalDependencies/UD_Telugu-MTG) (UD v2.17).

## Evaluation

Evaluation was performed on the UD_Telugu-MTG test split using the standard [CoNLL 2018 eval script](https://universaldependencies.org/conll18/conll18_ud_eval.py).

Two evaluation rows are reported:
- **Full-text (F1)**: raw text is segmented by [combo-seg](https://pypi.org/project/combo-seg/), then parsed and compared against gold — measures end-to-end pipeline performance including segmentation quality.
- **Aligned accuracy**: accuracy on correctly segmented (aligned) tokens — measures parsing quality on tokens that were correctly identified by the segmenter.

### Morphosyntactic Tagging

| Metric | Tokens | Sentences | Words | UPOS | XPOS | UFeats | AllTags | Lemmas |
| ------ | ------ | --------- | ----- | ---- | ---- | ------ | ------- | ------ |
| Full-text (F1) | 99.79 | 97.58 | 99.79 | 94.80 | 94.80 | 98.96 | 94.53 | 99.79 |
| Aligned accuracy | 0.00 | 0.00 | 0.00 | 95.00 | 95.00 | 99.17 | 94.72 | 100.00 |

### Dependency Parsing

| Metric | UAS | LAS | CLAS | MLAS | BLEX |
| ------ | --- | --- | ---- | ---- | ---- |
| Full-text (F1) | 90.64 | 83.16 | 79.43 | 76.19 | 79.43 |
| Aligned accuracy | 90.83 | 83.33 | 79.58 | 76.34 | 79.58 |


## Usage

Install the library from PyPI (assuming you have a virtual environment created):

```bash
pip install combo-nlp
```

The [combo-seg](https://pypi.org/project/combo-seg/) segmenter (used to split and
tokenise raw text) is installed automatically as a dependency of `combo-nlp`, so
no extra install step is needed.

```python
from combo import COMBO

# Load a pre-trained model with the corresponding combo-seg segmenter
nlp = COMBO("Telugu")

# Parse raw text (handles sentence splitting + tokenization)
result = nlp("వేగవంతమైన గోధుమ రంగు నక్క సోమరి కుక్క మీదుగా దూకుతుంది.")

# Inspect results
for sentence in result:
    for token in sentence:
        print(f"{token.form:<15} {token.lemma:<15} {token.upos:<8} head={token.head}  {token.deprel}")
```

Refer to the combo-nlp documentation for installation and usage instructions:

- [https://pypi.org/project/combo-nlp/](https://pypi.org/project/combo-nlp/)
- [https://pypi.org/project/combo-seg/](https://pypi.org/project/combo-seg/)

## License

The training data license: cc-by-sa-4.0 is derived from the Universal Dependencies treebank. For the full license terms of each treebank, please refer to the corresponding `LICENSE.txt` file in the treebank repository:

- [UD_Telugu-MTG LICENSE.txt](https://github.com/UniversalDependencies/UD_Telugu-MTG/blob/master/LICENSE.txt)


## Citation

If you use this model, please cite:

Ulewicz, M., Jabłońska, M., Klimaszewski, M., Przybyła, P., Pszenny, Ł., Rybak, P., Wiącek, M., & Wróblewska, A. (2026). *COMBO-NLP Models Trained on UD v2.17*. Zenodo. https://doi.org/10.5281/zenodo.19650523

```bibtex
@software{combo_nlp_2026,
  author    = {Ulewicz, Michał and Jabłońska, Maja and Klimaszewski, Mateusz and Przybyła, Piotr and Pszenny, Łukasz and Rybak, Piotr and Wiącek, Martyna and Wróblewska, Alina},
  title     = {{COMBO-NLP} Models Trained on {UD} v2.17},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.19650523},
  url       = {https://doi.org/10.5281/zenodo.19650523}
}
```



## Resources

- combo-nlp: [https://pypi.org/project/combo-nlp/](https://pypi.org/project/combo-nlp/)
- combo-seg: [https://pypi.org/project/combo-seg/](https://pypi.org/project/combo-seg/)
- UD_Telugu-MTG: [https://github.com/UniversalDependencies/UD_Telugu-MTG](https://github.com/UniversalDependencies/UD_Telugu-MTG)