Sentence Similarity
sentence-transformers
Safetensors
Luxembourgish
xlm-roberta
dataset_size:40000
loss:MSELoss
multilingual
Eval Results (legacy)
text-embeddings-inference
Instructions to use impresso-project/histlux-paraphrase-multilingual-mpnet-base-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use impresso-project/histlux-paraphrase-multilingual-mpnet-base-v2 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("impresso-project/histlux-paraphrase-multilingual-mpnet-base-v2") sentences = [ "Who is filming along?", "Wién filmt mat?", "Weider huet den Tatarescu drop higewisen, datt Rumänien durch seng krichsbedélegong op de 6eite vun den allie'erten 110.000 mann verluer hätt.", "Brambilla 130.08.03 St." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
Download tokenizer_config.json from impresso-project/histlux-paraphrase-multilingual-mpnet-base-v2: direct link, hf CLI and curl.
- Browser
- Download file 1.37 kB
-
https://huggingface.co/impresso-project/histlux-paraphrase-multilingual-mpnet-base-v2/resolve/main/tokenizer_config.json
- Command line
-
hf download hf://impresso-project/histlux-paraphrase-multilingual-mpnet-base-v2/tokenizer_config.json
-
curl -L -o tokenizer_config.json https://huggingface.co/impresso-project/histlux-paraphrase-multilingual-mpnet-base-v2/resolve/main/tokenizer_config.json
1.37 kB
| { | |
| "added_tokens_decoder": { | |
| "0": { | |
| "content": "<s>", | |
| "lstrip": false, | |
| "normalized": false, | |
| "rstrip": false, | |
| "single_word": false, | |
| "special": true | |
| }, | |
| "1": { | |
| "content": "<pad>", | |
| "lstrip": false, | |
| "normalized": false, | |
| "rstrip": false, | |
| "single_word": false, | |
| "special": true | |
| }, | |
| "2": { | |
| "content": "</s>", | |
| "lstrip": false, | |
| "normalized": false, | |
| "rstrip": false, | |
| "single_word": false, | |
| "special": true | |
| }, | |
| "3": { | |
| "content": "<unk>", | |
| "lstrip": false, | |
| "normalized": false, | |
| "rstrip": false, | |
| "single_word": false, | |
| "special": true | |
| }, | |
| "250001": { | |
| "content": "<mask>", | |
| "lstrip": true, | |
| "normalized": false, | |
| "rstrip": false, | |
| "single_word": false, | |
| "special": true | |
| } | |
| }, | |
| "bos_token": "<s>", | |
| "clean_up_tokenization_spaces": false, | |
| "cls_token": "<s>", | |
| "eos_token": "</s>", | |
| "extra_special_tokens": {}, | |
| "mask_token": "<mask>", | |
| "max_length": 128, | |
| "model_max_length": 128, | |
| "pad_to_multiple_of": null, | |
| "pad_token": "<pad>", | |
| "pad_token_type_id": 0, | |
| "padding_side": "right", | |
| "sep_token": "</s>", | |
| "stride": 0, | |
| "tokenizer_class": "XLMRobertaTokenizerFast", | |
| "truncation_side": "right", | |
| "truncation_strategy": "longest_first", | |
| "unk_token": "<unk>" | |
| } | |