Sentence Similarity
sentence-transformers
Safetensors
English
bert
feature-extraction
loss:SoftmaxLoss
Eval Results (legacy)
text-embeddings-inference
Instructions to use tomaarsen/bert-base-uncased-nli-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use tomaarsen/bert-base-uncased-nli-v1 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("tomaarsen/bert-base-uncased-nli-v1") sentences = [ "the guy is dead", "The dog is dead.", "The man is training the dog.", "People gather for an event." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
metadata
language:
- en
library_name: sentence-transformers
tags:
- sentence-transformers
- sentence-similarity
- feature-extraction
- loss:SoftmaxLoss
base_model: google-bert/bert-base-uncased
metrics:
- pearson_cosine
- spearman_cosine
- pearson_manhattan
- spearman_manhattan
- pearson_euclidean
- spearman_euclidean
- pearson_dot
- spearman_dot
- pearson_max
- spearman_max
widget:
- source_sentence: the guy is dead
sentences:
- The dog is dead.
- The man is training the dog.
- People gather for an event.
- source_sentence: the boy is five
sentences:
- The girl is five years old.
- A man sits in a hotel lobby.
- The man is laying on the couch.
- source_sentence: a guy is waxing
sentences:
- A woman is making music.
- A girl is laying in the pool
- She is the boy's aunt.
- source_sentence: Dog herding cows
sentences:
- A woman is walking her dog.
- Both people are standing up.
- The women are friends.
- source_sentence: There is a party
sentences:
- people take pictures
- A man is repainting a garage
- the crew all ate lunch alone
pipeline_tag: sentence-similarity
co2_eq_emissions:
emissions: 3.4540412355858656
energy_consumed: 0.008886090721390334
source: codecarbon
training_type: fine-tuning
on_cloud: false
cpu_model: 13th Gen Intel(R) Core(TM) i7-13700K
ram_total_size: 31.777088165283203
hours_used: 0.049
hardware_used: 1 x NVIDIA GeForce RTX 3090
model-index:
- name: SentenceTransformer based on google-bert/bert-base-uncased
results:
- task:
type: semantic-similarity
name: Semantic Similarity
dataset:
name: sts dev
type: sts-dev
metrics:
- type: pearson_cosine
value: 0.5998264726332272
name: Pearson Cosine
- type: spearman_cosine
value: 0.6439392261876368
name: Spearman Cosine
- type: pearson_manhattan
value: 0.6232915971361167
name: Pearson Manhattan
- type: spearman_manhattan
value: 0.6407370027700541
name: Spearman Manhattan
- type: pearson_euclidean
value: 0.6204725584722414
name: Pearson Euclidean
- type: spearman_euclidean
value: 0.6394239914170929
name: Spearman Euclidean
- type: pearson_dot
value: 0.4799617911944018
name: Pearson Dot
- type: spearman_dot
value: 0.4939854901099171
name: Spearman Dot
- type: pearson_max
value: 0.6232915971361167
name: Pearson Max
- type: spearman_max
value: 0.6439392261876368
name: Spearman Max
- task:
type: semantic-similarity
name: Semantic Similarity
dataset:
name: sts test
type: sts-test
metrics:
- type: pearson_cosine
value: 0.5516604742812986
name: Pearson Cosine
- type: spearman_cosine
value: 0.5840596347673308
name: Spearman Cosine
- type: pearson_manhattan
value: 0.5842488902993314
name: Pearson Manhattan
- type: spearman_manhattan
value: 0.5886614741524346
name: Spearman Manhattan
- type: pearson_euclidean
value: 0.582443715857982
name: Pearson Euclidean
- type: spearman_euclidean
value: 0.5869827075201962
name: Spearman Euclidean
- type: pearson_dot
value: 0.4054565422297012
name: Pearson Dot
- type: spearman_dot
value: 0.40476618101346834
name: Spearman Dot
- type: pearson_max
value: 0.5842488902993314
name: Pearson Max
- type: spearman_max
value: 0.5886614741524346
name: Spearman Max
SentenceTransformer based on google-bert/bert-base-uncased
This is a sentence-transformers model finetuned from google-bert/bert-base-uncased on the sentence-transformers/all-nli dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: google-bert/bert-base-uncased
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 768 tokens
- Similarity Function: Cosine Similarity
- Training Dataset:
- Language: en
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)
Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("tomaarsen/bert-base-uncased-nli-v1")
# Run inference
sentences = [
'There is a party',
'people take pictures',
'A man is repainting a garage',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings)
print(similarities.shape)
# [3, 3]
Evaluation
Metrics
Semantic Similarity
- Dataset:
sts-dev - Evaluated with
EmbeddingSimilarityEvaluator
| Metric | Value |
|---|---|
| pearson_cosine | 0.5998 |
| spearman_cosine | 0.6439 |
| pearson_manhattan | 0.6233 |
| spearman_manhattan | 0.6407 |
| pearson_euclidean | 0.6205 |
| spearman_euclidean | 0.6394 |
| pearson_dot | 0.48 |
| spearman_dot | 0.494 |
| pearson_max | 0.6233 |
| spearman_max | 0.6439 |
Semantic Similarity
- Dataset:
sts-test - Evaluated with
EmbeddingSimilarityEvaluator
| Metric | Value |
|---|---|
| pearson_cosine | 0.5517 |
| spearman_cosine | 0.5841 |
| pearson_manhattan | 0.5842 |
| spearman_manhattan | 0.5887 |
| pearson_euclidean | 0.5824 |
| spearman_euclidean | 0.587 |
| pearson_dot | 0.4055 |
| spearman_dot | 0.4048 |
| pearson_max | 0.5842 |
| spearman_max | 0.5887 |
Training Details
Training Dataset
sentence-transformers/all-nli
- Dataset: sentence-transformers/all-nli at cc6c526
- Size: 10,000 training samples
- Columns:
premise,hypothesis, andlabel - Approximate statistics based on the first 1000 samples:
premise hypothesis label type string string int details - min: 6 tokens
- mean: 17.38 tokens
- max: 52 tokens
- min: 4 tokens
- mean: 10.7 tokens
- max: 31 tokens
- 0: ~33.40%
- 1: ~33.30%
- 2: ~33.30%
- Samples:
premise hypothesis label A person on a horse jumps over a broken down airplane.A person is training his horse for a competition.1A person on a horse jumps over a broken down airplane.A person is at a diner, ordering an omelette.2A person on a horse jumps over a broken down airplane.A person is outdoors, on a horse.0 - Loss:
SoftmaxLoss
Evaluation Dataset
sentence-transformers/all-nli
- Dataset: sentence-transformers/all-nli at cc6c526
- Size: 1,000 evaluation samples
- Columns:
premise,hypothesis, andlabel - Approximate statistics based on the first 1000 samples:
premise hypothesis label type string string int details - min: 6 tokens
- mean: 18.44 tokens
- max: 57 tokens
- min: 5 tokens
- mean: 10.57 tokens
- max: 25 tokens
- 0: ~33.10%
- 1: ~33.30%
- 2: ~33.60%
- Samples:
premise hypothesis label Two women are embracing while holding to go packages.The sisters are hugging goodbye while holding to go packages after just eating lunch.1Two women are embracing while holding to go packages.Two woman are holding packages.0Two women are embracing while holding to go packages.The men are fighting outside a deli.2 - Loss:
SoftmaxLoss
Training Hyperparameters
Non-Default Hyperparameters
eval_strategy: stepsper_device_train_batch_size: 16per_device_eval_batch_size: 16num_train_epochs: 1warmup_ratio: 0.1bf16: True
All Hyperparameters
Click to expand
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Falseper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 1max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Nonedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: proportional
Training Logs
| Epoch | Step | Training Loss | loss | sts-dev_spearman_cosine | sts-test_spearman_cosine |
|---|---|---|---|---|---|
| 0 | 0 | - | - | 0.5931 | - |
| 0.16 | 100 | 1.056 | 0.9278 | 0.6555 | - |
| 0.32 | 200 | 0.8966 | 0.8751 | 0.6381 | - |
| 0.48 | 300 | 0.8646 | 0.8393 | 0.6170 | - |
| 0.64 | 400 | 0.8328 | 0.8100 | 0.5804 | - |
| 0.8 | 500 | 0.8307 | 0.7940 | 0.6413 | - |
| 0.96 | 600 | 0.8373 | 0.7602 | 0.6439 | - |
| 1.0 | 625 | - | - | - | 0.5841 |
Environmental Impact
Carbon emissions were measured using CodeCarbon.
- Energy Consumed: 0.009 kWh
- Carbon Emitted: 0.003 kg of CO2
- Hours Used: 0.049 hours
Training Hardware
- On Cloud: No
- GPU Model: 1 x NVIDIA GeForce RTX 3090
- CPU Model: 13th Gen Intel(R) Core(TM) i7-13700K
- RAM Size: 31.78 GB
Framework Versions
- Python: 3.11.6
- Sentence Transformers: 3.0.0.dev0
- Transformers: 4.41.0.dev0
- PyTorch: 2.3.0+cu121
- Accelerate: 0.26.1
- Datasets: 2.18.0
- Tokenizers: 0.19.1
Citation
BibTeX
Sentence Transformers and SoftmaxLoss
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}