--- language: - en license: apache-2.0 tags: - sentence-transformers - multi-vector - colbert - late-interaction - generated_from_trainer - dataset_size:640000 - loss:CachedMultiVectorMultipleNegativesRankingLoss widget: - text: "This confirmed our hypothesis that the predictor variables would be different\ \ for men and women. Our results are similar to those of Knechtle et al. [23]\ \ who reported that, for male ironman athletes, anthropometric variables were\ \ important, as percent body fat was significantly associated with total race\ \ time. In female triathletes, training volume showed a relationship to total\ \ race time, in corroboration of our study.\n\n Another interesting finding was\ \ that the coefficient of determination of the models was higher in women (r 2\ \ = 0.83) than in men (r 2 = 0.44). For women, the predicted race time did not\ \ correlate significantly (cor = 0.82, p-value = 0.09) with the achieved race\ \ time. For men, the predicted race time correlated significantly (cor = 0.84,\ \ p-value = 0.03) with the achieved race time. The differences in the coefficients\ \ of determination in the models might be explained by differences in anthropometric,\ \ training, and experience characteristics between women and men.\n\n A first\ \ important finding was that the personal best times in 5 km, 10 km, and half-marathon\ \ were the best predictors for female ultra-marathon performance. In the multiple\ \ regression models, the personal best half-marathon race time was significantly\ \ related to the ultra-marathon race time. Overall, it seems previous experience\ \ racing and fast personal best times are very important for ultra-marathon performance.\ \ This corroborates the results of Knechtle et al. [12] who examined 19 females\ \ in a 100 km ultra-marathon and found that the PBT in a marathon showed the highest\ \ correlation coefficient. Studies in other endurance sports disciplines such\ \ as triathlon showed personal best times in Olympic distance races were predictive\ \ in women for performance at Ironman distance [26] .\n\n Personal best times\ \ in marathon and in 5 km were associated with ultra-marathon race time for men.\ \ This adds to the bulk of knowledge available for males indicating that previous\ \ marathon personal best times seem to be a strong and independent predictor variable\ \ for ultra-endurance running performance in 100 km [15] , 350 km multi-day races\ \ [17] , and 24 h runs [17] . Previous studies have also shown that the personal\ \ best time in shorter races was also a predictor for Ironman race time in recreational\ \ male athletes [23, 27] , and PBT, not anthropometry or training volume, was\ \ associated with total race time in a triple-iron triathlon [24] . These findings\ \ of PBTs and high speed of running during training predicting ultramarathon performance\ \ reiterate the importance of intensity in training for men racing ultra-marathons.\n\ \n A recent study examined females racing in a 100 km distance ultra-marathon\ \ [12] . They found no association of race time with years running. We corroborate\ \ the results of Knechtle et al. [12] , as the variable years running was not\ \ associated with race time for women in this 62 km race. Rae et al. [28] examined\ \ the interaction of aging and racing on ultra-endurance running performance.\ \ Rae et al. [28] found that that overall athletes (18 women, 176 men) took approximately\ \ four years to reach peak running speed for a 56 km ultra-endurance race. It\ \ seemed that, regardless of the age at which the runners completed their first\ \ race, a period of about four years was required for the manifestation of adaptations\ \ associated with peak running performance during this ultra-endurance event.\ \ In our study, the average years running for females was 8.84, and so they were\ \ past this initial four years of improvement.\n\n Years running were associated\ \ to race time for males in bivariate analysis. This corroborates the results\ \ of Rae et al. [28] who studied mostly men (176 males, 18 females) and examined\ \ a similar distance to WUU2K (56 km versus 62 km of WUU2K). These findings contrast\ \ those of Knechtle et al. [17] , who examined multi-day racing male mountain\ \ runners, and of Knechtle et al. [15] , who examined 24 h race runners. For both\ \ of these studies, years running were not associated with ultra-marathon race\ \ time. Also, years running were not associated with marathon time for male marathoners\ \ [29] . Years running would seem to be more important for shorter runs, and maybe\ \ this would have to do with the intensity of training for performance in shorter\ \ runs, relying on less volume but more intense training, which would be easier\ \ for a non-novice runner.\n\n Recent studies show that age was an important predictor\ \ variable in ultra-marathon running [10] . Women's age was not significantly\ \ associated with race time in this 62 km race. This is in contrast to the results\ \ of Knechtle et al." - text: "Being a resource and role model for their colleagues, nurse champions can\ \ contribute to improved quality of palliative care, when they have sufficient\ \ clinical experience, improved knowledge of palliative care, improved teaching\ \ capacities, and acquired authority towards managers and colleagues. [25, 28]\ \ . Still, rigorous evaluation of the effects of nurse champions on the outcome\ \ of care is necessary. In this article we describe the study protocol of the\ \ PalTeC-H project: a study on understanding and improving Palliative and Terminal\ \ Care in the Hospital by implementing a palliative care network of nurse champions.\n\ \n \n\n Objectives of this study are (1) to explore and understand the impact\ \ of the quality of care on the quality of life at the end of life and the quality\ \ of dying in a hospital and (2) to investigate the contribution of a quality\ \ improvement intervention which consists of the implementation of a network of\ \ palliative care nurse champions. We define end-of-life care as care provided\ \ during the last three days of life (at most). We hypothesize the implementation\ \ of the network to result in more attention for palliative care, in improved\ \ and timely recognition of patients' palliative care needs, in more involvement\ \ of palliative care experts and, eventually, in improved quality of life during\ \ the last three days of life, improved quality of dying and increased satisfaction\ \ of bereaved relatives.\n\n The intervention consists of the establishment of\ \ a palliative care network of nurse champions which indirectly affects care by\ \ three main components: education, knowledge dissemination and support, plus\ \ several organizational elements (Table 1) . On intervention wards two staff\ \ nurses are appointed to be palliative care nurse championsfurther referred to\ \ as champions. Together they form the palliative care network coordinated by\ \ the multidisciplinary consultation team for pain and palliative care. Champions\ \ participate in monthly educational meetings of the network and in a targeted\ \ education programme of two days annually. The education programme includes palliative\ \ care knowledge and skills as well as organizational knowledge and skills, e.g.\ \ on planning dissemination of knowledge, in order to teach the champions to be\ \ an ambassador of palliative care on the wards and a role model for their colleagues.\ \ The educational strategy is based on the principles of constructivist learning\ \ and includes multiple approaches [37] . A senior nurse consultant, member of\ \ the multidisciplinary consultation team, is assigned to be the network coordinator,\ \ supported by the medical oncologist of the team. This network coordinator facilitates\ \ the learning process of champions, by organizing meetings and education programmes,\ \ and supporting champions individually in their development and in performing\ \ activities. The monthly meetings stimulate the incremental grow of knowledge.\ \ Working and learning in a network throughout the hospital give champions the\ \ opportunity to share knowledge and learn from others' experiences, and to capture\ \ knowledge from outside their own working environment [23, 28, 34] .\n\n Champions\ \ need to identify gaps in knowledge on and quality of palliative care on their\ \ ward and to raise health care givers' awareness on patients' palliative care\ \ needs. They have to organize educational activities, implement protocols on\ \ palliative and terminal care, and evaluate these activities at the end of each\ \ year.\n\n Assuming that 14 champions each spend eight hours per month on network\ \ activities, and that the coordinator spends 24 hours per month, the intervention\ \ costs are estimated at € 50.000 per year.\n\n All wards in a large general university\ \ hospital in the Netherlands participate in this study, including a specialized\ \ unit for palliative cancer care, but excluding the department of psychiatry\ \ and the Intensive Care departments.\n\n We collect data on adult patients who\ \ died at one of the 18 participating wards after having been admitted at least\ \ 6 hours prior to death.\n\n We designed a controlled before and after study\ \ with three phases: 1) pre-intervention phase (16 months); 2) phase in which\ \ the intervention is introduced (5 months); and 3) post-intervention phase (16\ \ months). The intervention, i.e. the appointment of two champions joining the\ \ network, is introduced in seven wards that regularly admit cancer patients or\ \ patients with other chronic and life threatening diseases, such as chronic cardiac\ \ diseases and COPD. Although there is not much evidence on the time needed to\ \ effectively disseminate expertise and knowledge into clinical practice [31]\ \ [32] [33] , we decided that the introduction phase lasts five months, as a run-up\ \ period to generate gradual changes in champions' behavior [16, 38] . In the\ \ 11 wards where the intervention is not introduced, the same measurements are\ \ performed to control for changes that are not due to the intervention, for example\ \ changes in hospital policy (Table 2) ." - text: "I n medical care, treatment decisions made by clinicians and patients are\ \ generally based-implicitly or explicitly-on predictions of comparative outcome\ \ risks under alternative treatment conditions. Randomized controlled trials (RCTs),\ \ widely accepted as the gold standard for determining causal effects, have provided\ \ the primary evidence for these predictions. However, there is mounting recognition\ \ within evidence-based medicine of the limitations of RCTs as tools to guide\ \ clinical decision making at the individual patient level (1) (2) (3) (4) . Although\ \ historically the overall summary result from randomized trials (\"average treatment\ \ effect\") has been the cornerstone of evidence-based clinical decisions, interest\ \ is growing in understanding how a treatment's effect can vary across patients-a\ \ concept described as heterogeneity of treatment effects (HTE) (5) (6) (7) (8)\ \ (9) (10) (11) .\n\n Much literature exists on the limitations of conventional\ \ \"1-variable-at-a-time\" subgroup analyses, which serially divide the trial\ \ population into groups (for example, male vs. female or old vs. young) and examine\ \ the contrast in the treatment effect across these groups (12) (13) (14) (15)\ \ (16) (17) (18) (19) (20) (21) (22) . The limitations include risks for false-negative\ \ and false-positive results due to low power for statistical interactions, weak\ \ prior theory on potential effect modifiers, and multiplicity (4, 10, (23) (24)\ \ (25) . These analyses are also incongruent with the way clinical decision making\ \ occurs at the level of the individual patient, because patients have multiple\ \ attributes simultaneously that can affect the tradeoffs between the benefits\ \ and harms of the intervention. Individual patients thus belong to multiple subgroups,\ \ each of which may yield a different estimate of the treatment effect (4, 10)\ \ .\n\n The PATH (Predictive Approaches to Treatment effect Heterogeneity) Statement\ \ offers guidance relevant for \"predictive\" approaches to HTE analysis (26)\ \ that are designed to address some of the limitations mentioned in the previous\ \ paragraph. The goal of predictive HTE analysis is to provide individualized\ \ predictions of treatment effect, specifically defined by the difference between\ \ expected potential outcomes of interest with one intervention versus an alternative\ \ (4, 8) . We refer to this as the \"individualized treatment effect.\" We avoid\ \ the term \"individual treatment effects\" because this latter term confusingly\ \ suggests that treatment effects can be estimated at the person level; such effects\ \ are inherently unobservable in parallel-group clinical trials because only 1\ \ of 2 counterfactual potential outcomes can be observed (10, 27) . Individualized\ \ treatment effects have also been termed \"conditional average treatment effects\"\ \ (28) , denoting that they are the averaged treatment effect in a subpopulation\ \ (that is, conditioned on a set of covariates). However, for prediction, we are\ \ specifically interested in identifying the best conditional average treatment\ \ effect given all available patient characteristics, where \"best\" is defined\ \ as that which best discriminates between future patients who do and do not benefit\ \ from a treatment to optimize decision making for individual patients (29). By\ \ accounting for multiple variables simultaneously, predictive HTE analysis is\ \ foundational to the concept of personalization in evidence-based medicine (4)\ \ . Statement guidance focuses on identifying \"clinically important HTE\" (4,\ \ 7, 10) , or variation in the risk difference across patient subgroups that may\ \ be sufficient to span important decision thresholds that reflect treatment-related\ \ harms and burdens. The statement offers guidance on 2 distinct approaches to\ \ predictive HTE analysis (4) . With a \"risk-modeling\" approach, a multivariable\ \ model that predicts risk for an outcome (usually the primary study outcome)\ \ is first identified from external sources (an \"external model\") or developed\ \ directly on the trial population without a term for treatment assignment (an\ \ \"internal model\"). This prediction model is then applied to disaggregate patients\ \ within trials to examine risk-based variation in treatment effects. In a second\ \ approach, \"effect modeling,\" a model is developed on RCT data with inclusion\ \ of a treatment assignment variable and potential inclusion of treatment interaction\ \ terms. These more flexible effect-modeling approaches have the potential to\ \ improve discrimination of patients who do and do not benefit, but they are especially\ \ vulnerable to overfitting and false discovery of promising subgroup effects\ \ (or they require very large databases that are well powered for the detection\ \ of interaction effects) (30). Both approaches can be used to predict individualized\ \ treatment effects-that is, the difference in expected outcome risks under 2\ \ alternative treatments, conditional on important clinical variables. A fuller\ \ introduction to risk and effect modeling is presented in prior literature (4)\ \ .\n\n In this PATH Statement explanation and elaboration, we expand on the intent\ \ and motivation (and reservations) regarding the statements, criteria, considerations,\ \ and caveats." - text: "Myocardial infarction (MI) remains the most frequent cardiovascular condition\ \ and can-beyond its immediate lethality-lead to cardiac failure and its associated\ \ late lethality. Cardiac failure is determined by the amount of myocardial tissue\ \ lost during ischaemia and, if reperfusion is achieved, the ensuing reperfusion\ \ injury, as well as by subsequent ventricular remodelling that adversely affects\ \ ventricular geometry.\n\n Ischaemic myocardial damage depends on cardiomyocyte\ \ apoptosis and necrosis. 1 Reperfusion injury is due to leucocyte-mediated cardiomyocyte\ \ bystander death during removal of necrosis. Growth of the defect also occurs\ \ secondary to stretch-induced tissue loss called non-ischaemic infarct expansion.\ \ 2, 3 Infarct healing can be divided into an early inflammatory and a late post-inflammatory\ \ phase. 4 The early inflammatory phase entails invasion of the infarcted tissue\ \ by leucocytes and removal of necrosis, population of the infarcted tissue by\ \ myofibroblasts and macrophages, and replacement of † These authors contributed\ \ equally to this work.\n\n necrosis by granulation tissue. The cellular events\ \ during this process in principle resemble wound healing in other tissues. 5\ \ However, haemodynamic strain and the release of pro-hypertrophic growth factors\ \ in the myocardium during healing are responsible for an additional phenomenon:\ \ adverse remodelling that determines longterm functional outcome after infarction.\ \ 6 Syndecans are a family of transmembrane heparansulfate proteoglycans that\ \ regulate cell -cell and cell -matrix interactions. 7 Increased levels of syndecan-4\ \ (Sdc4) were detected in the plasma of MI patients. 8 Expression of syndecan-1\ \ (Sdc1) and Sdc4 is increased in the infarcted and the remote myocardium in animal\ \ models of MI. 9 Sdc1, which is mainly expressed in inflammatory and vascular\ \ cells, has recently been shown to affect ischaemic myocardial damage by reducing\ \ inflammation and thereby left ventricular (LV) dilatation after ischaemia. 10\ \ Sdc4 is located within costamers and the Z-disc of cardiomyocytes, 11 which\ \ are thought to be important sites for mechano-sensing in cardiomyocytes. 12\ \ Sdc4 has been shown to translate mechanical stretch into cytoplasmic signalling\ \ in fibroblasts. 13 Wound healing in the skin is disturbed and delayed in Sdc4-deficient\ \ mice, 14 but the role of Sdc4 for myocardial wound healing and early remodelling\ \ has not been elucidated. We therefore examined the effects of Sdc4 deficiency\ \ on myocardial damage and early infarct healing in mouse models of myocardial\ \ ischaemia and infarction.\n\n \n\n This study was approved by the Institutional\ \ Review Board and performed in accordance with the Guide for the Care and Use\ \ of Laboratory Animals published by the US National Institutes of Health. Sdc4\ \ -/ -mice were backcrossed for more than 10 generations on C57BL/6 mice and ageand\ \ sex-matched Sdc4 -/ -and Sdc4 +/+ [wild-type (WT)] offspring of heterozygous\ \ matings were used for all studies. For myocardial ischaemia with reperfusion\ \ (MI/R) injury, the left coronary artery (LAD) was ligated for 30 min followed\ \ by 24 h of reperfusion. Area at risk (AAR) and infarct size were determined\ \ by TTC/Coomassie staining as described. 15 Data are presented as the average\ \ per cent infarct size per AAR. MI was induced by permanent ligation of the LAD\ \ as published previously. 16 Hearts were taken out 7 days later for molecular\ \ and histological analyses. Blood was collected from the retrobulbar plexus 24\ \ h after MI or was drawn from the inferior vena cava before mice were sacrificed.\ \ Plasma was separated and troponin T levels were assessed by using Elecsys Troponin\ \ T high-sensitive test (Roche Diagnostics, Mannheim, Germany). 17 \n\n Primary\ \ neonatal rat ventricular cardiomyocytes were harvested from 1-to 3-day-old Sprague-Dawley\ \ rat pups (Charles River, Sulzfeld, Germany) as published elsewhere. 18 Cardiomyocytes\ \ were transfected with siRNA to Sdc4 (151941, Applied Biosystems, Darmstadt,\ \ Germany) or nonspecific (scr)-siRNA (1027281, Qiagen, Hilden, Germany) using\ \ RNAiMAX (Invitrogen, Karlsruhe, Germany)." - text: "ACR20 response rates were significantly higher with CZP plus MTX than placebo\ \ plus MTX at Week 1 (22.9 and 14.3% with CZP 200 mg plus MTX vs 5.6 and 3.3%\ \ with placebo plus MTX in the RAPID 1 and 2 trials, respectively) [4, 5] . ACR20\ \ response rates peaked at Week 12 in both studies (63.8 and 62.7% for CZP 200\ \ mg vs 18.3 and 12.7% for placebo in RAPID 1 and 2, respectively; both P < 0.001).\ \ At Week 24, ACR20 response rates were 58.8 and 57.3% for patients receiving\ \ CZP 200 mg plus MTX, respectively, vs 13.6 and 8.7%. The ITT populations for\ \ RAPID 1 and 2 consisted of all patients who were randomized into the studies;\ \ the modified ITT population for FAST4WARD consisted of all randomized patients\ \ who had taken one or more dose of study medication. Adapted from Mease [21]\ \ with permission of Future Medicine Ltd. CV: coefficient of variation; ITT: intention-to-treat;\ \ NA: not applicable.\n\n Significantly higher ACR50 and ACR70 response rates\ \ for CZP vs placebo groups were seen from Weeks 2 and 4 in RAPID 1, and Weeks\ \ 6 and 20 in RAPID 2, respectively. Responses were sustained to the end of the\ \ trials (Week 52 in RAPID 1 and Week 24 in RAPID 2; Table 2 ), and were similar\ \ in the CZP 400 mg plus MTX groups. CZP treatment also yielded significant improvements\ \ in all ACR core component scores, including reductions in swollen and tender\ \ joint scores and improvements in both patient's and physician's global assessments\ \ of disease activity, by Week 1 that were sustained throughout both studies [4,\ \ 5] . Treatment with CZP plus MTX was associated with significantly greater improvements\ \ in disease activity from Week 1, as evidenced by DAS-28 (ESR) scores, throughout\ \ both trials (P < 0.001 at all time points) [4, 5] . At Week 1, mean change from\ \ baseline in DAS-28 was À0.8 with CZP 200 mg and À0.3 with placebo in RAPID 1,\ \ and À0.8 with CZP 200 mg and À0.2 with placebo in RAPID 2. Improvements were\ \ sustained to the end of both studies (52 or 24 weeks, respectively; Fig. 1 ),\ \ and were similar with the CZP 400 mg dose. In RAPID 2, DAS-28 remission was\ \ observed in 9.4% of patients treated with CZP 200 mg plus MTX compared with\ \ only 0.8% of patients in the placebo group [5] .\n\n Both trials investigated\ \ the effects of CZP on the progression of joint damage. In RAPID 1, the mean\ \ (S.D.) change in mTSS from baseline to Week 52, which was a co-primary endpoint\ \ of the study, was significantly lower in patients receiving CZP 200 mg plus\ \ MTX [0.4 (5.7) in the CZP 200 mg group] compared with patients receiving placebo\ \ plus MTX [2.8 (7.8); P < 0.001] [4] . The changes were also significantly lower\ \ in the CZP plus MTX groups vs the placebo plus MTX group at Week 24 (P < 0.001).\ \ At both time points, significantly lower mean changes from baseline in both\ \ erosion (Week 24: 0 vs 0.7, Week 52: 0.1 vs 1.5; P < 0. [5] . Patients in the\ \ CZP 200 mg group in RAPID 2 also had significantly lower erosion (mean change\ \ from baseline: 0.1 vs 0.7) and joint space narrowing (mean change from baseline:\ \ 0.1 vs 0.5) subscores (P 4 0.01). Results for patients receiving the 400-mg\ \ dose were similar. An analysis of joint damage in patients who withdrew from\ \ the trials at Week 16 due to ACR20 non-response at Weeks 12 and 14 (as mandated\ \ by the study protocol) found that radiographic progression was inhibited by\ \ CZP plus MTX despite the fact that these patients did not meet the threshold\ \ for a clinical response [4, 5] ." pipeline_tag: feature-extraction library_name: sentence-transformers metrics: - maxsim_accuracy@1 - maxsim_accuracy@3 - maxsim_accuracy@5 - maxsim_accuracy@10 - maxsim_precision@1 - maxsim_precision@3 - maxsim_precision@5 - maxsim_precision@10 - maxsim_recall@1 - maxsim_recall@3 - maxsim_recall@5 - maxsim_recall@10 - maxsim_ndcg@10 - maxsim_mrr@10 - maxsim_map@100 model-index: - name: ColBERT gte-modernbert-base trained on MIRIAD question-passage pairs results: - task: type: multi-vector-information-retrieval name: Multi Vector Information Retrieval dataset: name: miriad eval type: miriad_eval metrics: - type: maxsim_accuracy@1 value: 0.974 name: Maxsim Accuracy@1 - type: maxsim_accuracy@3 value: 0.99 name: Maxsim Accuracy@3 - type: maxsim_accuracy@5 value: 0.995 name: Maxsim Accuracy@5 - type: maxsim_accuracy@10 value: 0.997 name: Maxsim Accuracy@10 - type: maxsim_precision@1 value: 0.974 name: Maxsim Precision@1 - type: maxsim_precision@3 value: 0.32999999999999996 name: Maxsim Precision@3 - type: maxsim_precision@5 value: 0.19900000000000004 name: Maxsim Precision@5 - type: maxsim_precision@10 value: 0.09970000000000001 name: Maxsim Precision@10 - type: maxsim_recall@1 value: 0.974 name: Maxsim Recall@1 - type: maxsim_recall@3 value: 0.99 name: Maxsim Recall@3 - type: maxsim_recall@5 value: 0.995 name: Maxsim Recall@5 - type: maxsim_recall@10 value: 0.997 name: Maxsim Recall@10 - type: maxsim_ndcg@10 value: 0.9863523881462793 name: Maxsim Ndcg@10 - type: maxsim_mrr@10 value: 0.9828250000000002 name: Maxsim Mrr@10 - type: maxsim_map@100 value: 0.9830133116883116 name: Maxsim Map@100 --- # ColBERT gte-modernbert-base trained on MIRIAD question-passage pairs This is a [Multi-Vector Encoder](https://www.sbert.net/docs/multi_vector_encoder/usage/usage.html) model trained on the miriad-4.4_m-split dataset using the [sentence-transformers](https://www.SBERT.net) library. It maps inputs to sequences of 128-dimensional token-level vectors and scores them with late interaction (MaxSim), useful for semantic search with late interaction. ## Model Details ### Model Description - **Model Type:** Multi-Vector Encoder - **Maximum Sequence Length:** 8192 tokens - **Output Dimensionality:** 128 dimensions - **Similarity Function:** maxsim - **Supported Modality:** Text - **Training Dataset:** - miriad-4.4_m-split - **Language:** en - **License:** apache-2.0 ### Model Sources - **Documentation:** [Sentence Transformers Documentation](https://sbert.net) - **Documentation:** [Multi-Vector Encoder Documentation](https://www.sbert.net/docs/multi_vector_encoder/usage/usage.html) - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers) - **Hugging Face:** [Multi-Vector Encoders on Hugging Face](https://huggingface.co/models?library=sentence-transformers&other=multi-vector) ### Full Model Architecture ``` MultiVectorEncoder( (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'query_expansion': {'strategy': 'min', 'attend': False, 'token': None, 'length': 32}, 'architecture': 'ModernBertModel'}) (1): Dense({'in_features': 768, 'out_features': 128, 'bias': False, 'activation_function': 'torch.nn.modules.linear.Identity', 'module_input_name': 'token_embeddings', 'module_output_name': 'token_embeddings'}) (2): MultiVectorMask({'skiplist_words': ['!', '"', '#', '$', '%', '&', "'", '(', ')', '*', '+', ',', '-', '.', '/', ':', ';', '<', '=', '>', '?', '@', '[', '\\', ']', '^', '_', '`', '{', '|', '}', '~'], 'keep_only_token_ids': None}) (3): Normalize({'module_input_name': 'token_embeddings', 'module_output_name': 'token_embeddings'}) ) ``` ## Usage ### Direct Usage (Sentence Transformers) First install the Sentence Transformers library: ```bash pip install -U sentence-transformers ``` Then you can load this model and run inference. ```python from sentence_transformers import MultiVectorEncoder # Download from the 🤗 Hub model = MultiVectorEncoder("tomaarsen/multivector-gte-modernbert-base-miriad") # Run inference: each input becomes a sequence of per-token vectors (variable length). queries = [ 'How does treatment with CZP plus MTX impact disease activity in patients with rheumatoid arthritis?\n', ] documents = [ "ACR20 response rates were significantly higher with CZP plus MTX than placebo plus MTX at Week 1 (22.9 and 14.3% with CZP 200 mg plus MTX vs 5.6 and 3.3% with placebo plus MTX in the RAPID 1 and 2 trials, respectively) [4, 5] . ACR20 response rates peaked at Week 12 in both studies (63.8 and 62.7% for CZP 200 mg vs 18.3 and 12.7% for placebo in RAPID 1 and 2, respectively; both P < 0.001). At Week 24, ACR20 response rates were 58.8 and 57.3% for patients receiving CZP 200 mg plus MTX, respectively, vs 13.6 and 8.7%. The ITT populations for RAPID 1 and 2 consisted of all patients who were randomized into the studies; the modified ITT population for FAST4WARD consisted of all randomized patients who had taken one or more dose of study medication. Adapted from Mease [21] with permission of Future Medicine Ltd. CV: coefficient of variation; ITT: intention-to-treat; NA: not applicable.\n\n Significantly higher ACR50 and ACR70 response rates for CZP vs placebo groups were seen from Weeks 2 and 4 in RAPID 1, and Weeks 6 and 20 in RAPID 2, respectively. Responses were sustained to the end of the trials (Week 52 in RAPID 1 and Week 24 in RAPID 2; Table 2 ), and were similar in the CZP 400 mg plus MTX groups. CZP treatment also yielded significant improvements in all ACR core component scores, including reductions in swollen and tender joint scores and improvements in both patient's and physician's global assessments of disease activity, by Week 1 that were sustained throughout both studies [4, 5] . Treatment with CZP plus MTX was associated with significantly greater improvements in disease activity from Week 1, as evidenced by DAS-28 (ESR) scores, throughout both trials (P < 0.001 at all time points) [4, 5] . At Week 1, mean change from baseline in DAS-28 was À0.8 with CZP 200 mg and À0.3 with placebo in RAPID 1, and À0.8 with CZP 200 mg and À0.2 with placebo in RAPID 2. Improvements were sustained to the end of both studies (52 or 24 weeks, respectively; Fig. 1 ), and were similar with the CZP 400 mg dose. In RAPID 2, DAS-28 remission was observed in 9.4% of patients treated with CZP 200 mg plus MTX compared with only 0.8% of patients in the placebo group [5] .\n\n Both trials investigated the effects of CZP on the progression of joint damage. In RAPID 1, the mean (S.D.) change in mTSS from baseline to Week 52, which was a co-primary endpoint of the study, was significantly lower in patients receiving CZP 200 mg plus MTX [0.4 (5.7) in the CZP 200 mg group] compared with patients receiving placebo plus MTX [2.8 (7.8); P < 0.001] [4] . The changes were also significantly lower in the CZP plus MTX groups vs the placebo plus MTX group at Week 24 (P < 0.001). At both time points, significantly lower mean changes from baseline in both erosion (Week 24: 0 vs 0.7, Week 52: 0.1 vs 1.5; P < 0. [5] . Patients in the CZP 200 mg group in RAPID 2 also had significantly lower erosion (mean change from baseline: 0.1 vs 0.7) and joint space narrowing (mean change from baseline: 0.1 vs 0.5) subscores (P 4 0.01). Results for patients receiving the 400-mg dose were similar. An analysis of joint damage in patients who withdrew from the trials at Week 16 due to ACR20 non-response at Weeks 12 and 14 (as mandated by the study protocol) found that radiographic progression was inhibited by CZP plus MTX despite the fact that these patients did not meet the threshold for a clinical response [4, 5] .", 'Five minutes after atropine, the R:T ratio increased from 1.15 (0.4) to 1.40 (0.6) (P < 0.01); at 30 min it was 1.51 (0.7) (P < 0.001) and at 60 min it was 1.33 (0.5) (P < 0.05). The R-wave amplitude was not affected by atropine. No changes in heart rate, QTc interval, RSA and R:T ratio occurred after placebo. COMMENT These data show that, in the presence of vagal block by atropine, the QTc interval increased significantly and the T-wave of the ECG was flattened.\n\n We chose a relatively large dose of atropine to ensure parasympathetic block as confirmed by the disappearance of RSA. Day, McComp and Campbell fl] have suggested that QT dispersion (interlead variability) gives an indication of arrythmogenicity and repolarization. We used a single lead V 2 which, according to the same group, provides the closest approximation to maximum QT interval [4] . They also accept the validity of a single lead value for QTc when changes are monitored. The flattened T-wave after atropine in our volunteers probably also reflected irregularity in repolarization.\n\n Atropine has been shown to increase the incidence of cardiac arrhythmia during induction of anaesthesia [3] . In addition, i.v. atropine has been shown to cause ventricular tachycardia in a patient with a prolonged QT interval syndrome [5] . Inhibition of the sympathoadrenal tone by opioids shortens the QTc interval in patients with vagal block. Vagal stimulation protects the heart against arrhythmogenic vulnerability [2] and against prolongation of the QT interval. In our study, the QTc interval was prolonged, probably because sympathoadrenal tone became dominant after parasympathetic block by atropine.\n\n In diabetic patients, vagal denervation develops gradually. Maintenance of remaining borderline vagal function by avoiding anticholinergics may be of value in diabetic patients, as serious cardiac arrhythmia has been described in these patients during anaesthesia and after atropine. Furthermore, ventricular fibrillation after i.v. atropine for bradycardia has been shown to occur in acute myocardial infarction [6] . The routine use of anticholinergics at induction of anaesthesia must be seriously questioned.', 'This occurs since, in the folded state, the dansyl group is encapsulated in the hydrophobic cavity of the β-cyclodextrin ring resulting in a net fluorescence enhancement [99] . As a further development of this work, Riccardi and co-workers have described a tris-conjugated TBA 15 (tris-mTBA), equipped with a dansyl, a β-cyclodextrin and a biotin tag at the ends. This novel design has allowed the incorporation of TBA 15 onto streptavidin-coated NPs, leading to a remarkable increase of its anticoagulant properties. The developed systems have provided the basis for suitable aptamer-based devices for theranostic applications, allowing simultaneously both fluorescence-based detection and modulation of the thrombin activity [101] .\n\n Notably, in addition to the sensing approaches based on conformational switch random coil-G-quadruplex structure, also thrombin-induced changes starting from a hairpin structure are possible if the aptamer is properly engineered. In this context, Hamaguchi et al. have described a TBA 15 elongated at the 5 -end with few nucleotides complementary to the 3 -end and therefore able to adopt a stem-loop structure or hairpin [102] . In addition, the aptamer is equipped with a fluorescent/quencher pair, i.e., a fluorescein and a dabcyl moiety at the 5 -and 3 -end, respectively. In the absence of thrombin, the close proximity between the two reporter groups in the hairpin structure determines fluorescence quenching. After thrombin recognition, the stem-loop structure is destabilized in favour of interactions with the protein. Under these conditions, the fluorescent dye and the quencher are distant, thus allowing a "turn-on" of the fluorescence signal, indicative of the binding with the target molecule ( Figure 7c ).\n\n Alternative approaches for "structure switch signalling aptamers" are reported by Nutiu and Li [103] . Their strategy for designing aptamer-based fluorescent reporters involves structural switches from DNA/DNA duplex to DNA/target complex. In this study, the aptamer beacon consists of a tripartite duplex structure including a 5 -fluorescein-labeled oligomer (FDNA), a 3 -dabcyl-labeled oligomer (QDNA) and a longer oligonucleotide sequence comprising Stem-1 and Stem-2, complementary to FDNA and QDNA, respectively. Stem-2 also contains the TBA 15 sequence in a partial overhang (Figure 8a ). In the absence of the target protein, the aptamer naturally binds to FDNA and QDNA, bringing the fluorophore and the quencher in close proximity and thus completely inhibiting the fluorescence signal. The presence of thrombin triggers the formation of the aptamer-target complex, causing the release of QDNA and fully restoring the fluorescence emission.\n\n Cancers 2017, 9, 174 10 of 43\n\n Notably, in addition to the sensing approaches based on conformational switch random coil-G-quadruplex structure, also thrombin-induced changes starting from a hairpin structure are possible if the aptamer is properly engineered. In this context, Hamaguchi et al. have described a TBA15 elongated at the 5′-end with few nucleotides complementary to the 3′-end and therefore able to adopt a stem-loop structure or hairpin [102] . In addition, the aptamer is equipped with a fluorescent/quencher pair, i.e., a fluorescein and a dabcyl moiety at the 5′-and 3′-end, respectively. In the absence of thrombin, the close proximity between the two reporter groups in the hairpin structure determines fluorescence quenching. After thrombin recognition, the stem-loop structure is destabilized in favour of interactions with the protein. Under these conditions, the fluorescent dye and the quencher are distant, thus allowing a "turn-on" of the fluorescence signal, indicative of the binding with the target molecule ( Figure 7c ).\n\n Alternative approaches for "structure switch signalling aptamers" are reported by Nutiu and Li [103] . Their strategy for designing aptamer-based fluorescent reporters involves structural switches from DNA/DNA duplex to DNA/target complex. In this study, the aptamer beacon consists of a tripartite duplex structure including a 5′-fluorescein-labeled oligomer (FDNA), a 3′-dabcyl-labeled oligomer (QDNA) and a longer oligonucleotide sequence comprising Stem-1 and Stem-2, complementary to FDNA and QDNA, respectively. Stem-2 also contains the TBA15 sequence in a partial overhang (Figure 8a ).', ] query_embeddings = model.encode_query(queries) document_embeddings = model.encode_document(documents) print(query_embeddings[0].shape, document_embeddings[0].shape) # (32, 128) (814, 128) # Get the MaxSim similarity scores similarities = model.similarity(query_embeddings, document_embeddings) print(similarities) # tensor([[26.4624, 9.1034, 3.9515]]) ``` ## Evaluation ### Metrics #### Multi Vector Information Retrieval * Dataset: `miriad_eval` * Evaluated with [MultiVectorInformationRetrievalEvaluator](https://sbert.net/docs/package_reference/multi_vector_encoder/evaluation.html#sentence_transformers.multi_vector_encoder.evaluation.MultiVectorInformationRetrievalEvaluator) | Metric | Value | |:--------------------|:-----------| | maxsim_accuracy@1 | 0.974 | | maxsim_accuracy@3 | 0.99 | | maxsim_accuracy@5 | 0.995 | | maxsim_accuracy@10 | 0.997 | | maxsim_precision@1 | 0.974 | | maxsim_precision@3 | 0.33 | | maxsim_precision@5 | 0.199 | | maxsim_precision@10 | 0.0997 | | maxsim_recall@1 | 0.974 | | maxsim_recall@3 | 0.99 | | maxsim_recall@5 | 0.995 | | maxsim_recall@10 | 0.997 | | **maxsim_ndcg@10** | **0.9864** | | maxsim_mrr@10 | 0.9828 | | maxsim_map@100 | 0.983 | ## Training Details ### Training Dataset #### miriad-4.4_m-split * Dataset: miriad-4.4_m-split * Size: 640,000 training samples * Columns: question and passage_text * Approximate statistics based on the first 100 samples: | | question | passage_text | |:---------|:-----------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------| | type | string | string | | modality | text | text | | details | | | * Samples: | question | passage_text | |:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | What factors may contribute to increased pulmonary conduit durability in patients who undergo the Ross operation compared to those with right ventricular outflow tract obstruction?
| I n 1966, Ross and Somerville 1 reported the first use of an aortic homograft to establish right ventricle-to-pulmonary artery continuity in a patient with tetralogy of Fallot and pulmonary atresia. Since that time, pulmonary position homografts have been used in a variety of right-sided congenital heart lesions. Actuarial 5-year homograft survivals for cryopreserved homografts are reported to range between 55% and 94%, with the shortest durability noted in patients less than 2 years of age. 4 Pulmonary position homografts also are used to replace pulmonary autografts explanted to repair left-sided outflow disease (the Ross operation). Several factors may be likely to favor increased pulmonary conduit durability in Ross patients compared with those with right ventricular outflow tract obstruction, including later age at operation (allowing for larger homografts), more normal pulmonary artery architecture, absence of severe right ventricular hypertrophy, and more natural positioning of ... | | How does MCAM expression in hMSC affect the growth and maintenance of hematopoietic progenitors? | After culture in a 3-dimensional hydrogel-based matrix, which constitutes hypoxic conditions, MCAM expression is lost. Concordantly, Tormin et al. demonstrated that MCAM is down-regulated under hypoxic conditions. 10 Furthermore, it was shown by others and our group that oxygen tension causes selective modification of hematopoietic cell and mesenchymal stromal cell interactions in co-culture systems as well as influence HSPC metabolism. [44] [45] [46] Thus, the observed differences between Sharma et al. and our data in HSPC supporting capacity of hMSC are likely due to the different culture conditions used. Further studies are required to clarify the influence of hypoxia in our model system. Altogether these findings provide further evidence for the importance of MCAM in supporting HSPC. Furthermore, previous reports have shown that MCAM is down-regulated in MSC after several passages as well as during aging and differentiation. 19, 47 Interestingly, MCAM overexpression in hMSC enhance... | | What is the relationship between Fanconi anemia and breast and ovarian cancer susceptibility genes?
| ( 31 ) , of which 5% -10 % may be caused by genetic factors ( 32 ) , up to half a million of these patients may be at risk of secondary hereditary neoplasms. The historic observation of twofold to fi vefold increased risks of cancers of the ovary, thyroid, and connective tissue after breast cancer ( 33 ) presaged the later syndromic association of these tumors with inherited mutations of BRCA1, BRCA2, PTEN, and p53 ( 16 ) . By far the largest cumulative risk of a secondary cancer in BRCA mutation carriers is associated with cancer in the contralateral breast, which may reach a risk of 29.5% at 10 years ( 34 ) . The Breast Cancer Linkage Consortium ( 35 , 36 ) also documented threefold to fi vefold increased risks of subsequent cancers of prostate, pancreas, gallbladder, stomach, skin (melanoma), and uterus in BRCA2 mutation carriers and twofold increased risks of prostate and pancreas cancer in BRCA1 mutation carriers; these results are based largely on self-reported family history inf... | * Loss: [CachedMultiVectorMultipleNegativesRankingLoss](https://sbert.net/docs/package_reference/multi_vector_encoder/losses.html#cachedmultivectormultiplenegativesrankingloss) with these parameters: ```json { "score_metric": "colbert_scores", "mini_batch_size": 8, "mini_batch_num_tokens": null, "score_mini_batch_size": 8, "scale": 1.0, "size_average": true, "gather_across_devices": false } ``` ### Training Hyperparameters #### Non-Default Hyperparameters - `per_device_train_batch_size`: 128 - `num_train_epochs`: 1 - `learning_rate`: 3e-05 - `warmup_steps`: 0.05 - `bf16`: True - `load_best_model_at_end`: True - `batch_sampler`: no_duplicates - `max_length`: 1024 #### All Hyperparameters
Click to expand - `per_device_train_batch_size`: 128 - `num_train_epochs`: 1 - `max_steps`: -1 - `learning_rate`: 3e-05 - `lr_scheduler_type`: linear - `lr_scheduler_kwargs`: None - `warmup_steps`: 0.05 - `optim`: adamw_torch_fused - `optim_args`: None - `weight_decay`: 0.0 - `adam_beta1`: 0.9 - `adam_beta2`: 0.999 - `adam_epsilon`: 1e-08 - `optim_target_modules`: None - `gradient_accumulation_steps`: 1 - `average_tokens_across_devices`: True - `max_grad_norm`: 1.0 - `label_smoothing_factor`: 0.0 - `bf16`: True - `fp16`: False - `bf16_full_eval`: False - `fp16_full_eval`: False - `tf32`: None - `gradient_checkpointing`: False - `gradient_checkpointing_kwargs`: None - `torch_compile`: False - `torch_compile_backend`: None - `torch_compile_mode`: None - `use_liger_kernel`: False - `liger_kernel_config`: None - `use_cache`: False - `neftune_noise_alpha`: None - `torch_empty_cache_steps`: None - `auto_find_batch_size`: False - `log_on_each_node`: True - `logging_nan_inf_filter`: True - `include_num_input_tokens_seen`: no - `log_level`: passive - `log_level_replica`: warning - `disable_tqdm`: False - `project`: huggingface - `trackio_space_id`: None - `trackio_bucket_id`: None - `trackio_static_space_id`: None - `per_device_eval_batch_size`: 8 - `prediction_loss_only`: True - `eval_on_start`: False - `eval_do_concat_batches`: True - `eval_use_gather_object`: False - `eval_accumulation_steps`: None - `include_for_metrics`: [] - `batch_eval_metrics`: False - `save_only_model`: False - `save_on_each_node`: False - `enable_jit_checkpoint`: False - `push_to_hub`: False - `hub_private_repo`: None - `hub_model_id`: None - `hub_strategy`: every_save - `hub_always_push`: False - `hub_revision`: None - `load_best_model_at_end`: True - `ignore_data_skip`: False - `restore_callback_states_from_checkpoint`: False - `full_determinism`: False - `seed`: 42 - `data_seed`: None - `use_cpu`: False - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None} - `parallelism_config`: None - `dataloader_drop_last`: False - `dataloader_num_workers`: 0 - `dataloader_pin_memory`: True - `dataloader_persistent_workers`: False - `dataloader_prefetch_factor`: None - `remove_unused_columns`: True - `label_names`: None - `train_sampling_strategy`: random - `length_column_name`: length - `ddp_find_unused_parameters`: None - `ddp_bucket_cap_mb`: None - `ddp_broadcast_buffers`: False - `ddp_static_graph`: None - `ddp_backend`: None - `ddp_timeout`: 1800 - `fsdp`: None - `fsdp_config`: None - `deepspeed`: None - `debug`: [] - `skip_memory_metrics`: True - `do_predict`: False - `resume_from_checkpoint`: None - `warmup_ratio`: None - `local_rank`: -1 - `prompts`: None - `batch_sampler`: no_duplicates - `multi_dataset_batch_sampler`: proportional - `router_mapping`: {} - `learning_rate_mapping`: {} - `max_length`: 1024
### Training Logs
Click to expand | Epoch | Step | Training Loss | miriad_eval_maxsim_ndcg@10 | |:-----:|:----:|:-------------:|:--------------------------:| | -1 | -1 | - | 0.9176 | | 0.005 | 25 | 0.2384 | - | | 0.01 | 50 | 0.1516 | - | | 0.015 | 75 | 0.0738 | - | | 0.02 | 100 | 0.0545 | - | | 0.025 | 125 | 0.0356 | - | | 0.03 | 150 | 0.0273 | - | | 0.035 | 175 | 0.0282 | - | | 0.04 | 200 | 0.0200 | - | | 0.045 | 225 | 0.0218 | - | | 0.05 | 250 | 0.0117 | - | | 0.055 | 275 | 0.0121 | - | | 0.06 | 300 | 0.0120 | - | | 0.065 | 325 | 0.0137 | - | | 0.07 | 350 | 0.0125 | - | | 0.075 | 375 | 0.0154 | - | | 0.08 | 400 | 0.0123 | - | | 0.085 | 425 | 0.0100 | - | | 0.09 | 450 | 0.0112 | - | | 0.095 | 475 | 0.0109 | - | | 0.1 | 500 | 0.0092 | 0.9788 | | 0.105 | 525 | 0.0102 | - | | 0.11 | 550 | 0.0105 | - | | 0.115 | 575 | 0.0062 | - | | 0.12 | 600 | 0.0113 | - | | 0.125 | 625 | 0.0063 | - | | 0.13 | 650 | 0.0118 | - | | 0.135 | 675 | 0.0068 | - | | 0.14 | 700 | 0.0086 | - | | 0.145 | 725 | 0.0077 | - | | 0.15 | 750 | 0.0103 | - | | 0.155 | 775 | 0.0169 | - | | 0.16 | 800 | 0.0097 | - | | 0.165 | 825 | 0.0110 | - | | 0.17 | 850 | 0.0072 | - | | 0.175 | 875 | 0.0072 | - | | 0.18 | 900 | 0.0081 | - | | 0.185 | 925 | 0.0071 | - | | 0.19 | 950 | 0.0091 | - | | 0.195 | 975 | 0.0111 | - | | 0.2 | 1000 | 0.0072 | 0.9794 | | 0.205 | 1025 | 0.0068 | - | | 0.21 | 1050 | 0.0076 | - | | 0.215 | 1075 | 0.0076 | - | | 0.22 | 1100 | 0.0076 | - | | 0.225 | 1125 | 0.0146 | - | | 0.23 | 1150 | 0.0068 | - | | 0.235 | 1175 | 0.0062 | - | | 0.24 | 1200 | 0.0094 | - | | 0.245 | 1225 | 0.0063 | - | | 0.25 | 1250 | 0.0103 | - | | 0.255 | 1275 | 0.0070 | - | | 0.26 | 1300 | 0.0075 | - | | 0.265 | 1325 | 0.0072 | - | | 0.27 | 1350 | 0.0053 | - | | 0.275 | 1375 | 0.0043 | - | | 0.28 | 1400 | 0.0091 | - | | 0.285 | 1425 | 0.0092 | - | | 0.29 | 1450 | 0.0077 | - | | 0.295 | 1475 | 0.0092 | - | | 0.3 | 1500 | 0.0064 | 0.9766 | | 0.305 | 1525 | 0.0069 | - | | 0.31 | 1550 | 0.0069 | - | | 0.315 | 1575 | 0.0061 | - | | 0.32 | 1600 | 0.0070 | - | | 0.325 | 1625 | 0.0074 | - | | 0.33 | 1650 | 0.0059 | - | | 0.335 | 1675 | 0.0069 | - | | 0.34 | 1700 | 0.0071 | - | | 0.345 | 1725 | 0.0056 | - | | 0.35 | 1750 | 0.0082 | - | | 0.355 | 1775 | 0.0059 | - | | 0.36 | 1800 | 0.0059 | - | | 0.365 | 1825 | 0.0072 | - | | 0.37 | 1850 | 0.0073 | - | | 0.375 | 1875 | 0.0037 | - | | 0.38 | 1900 | 0.0072 | - | | 0.385 | 1925 | 0.0045 | - | | 0.39 | 1950 | 0.0055 | - | | 0.395 | 1975 | 0.0062 | - | | 0.4 | 2000 | 0.0059 | 0.9781 | | 0.405 | 2025 | 0.0057 | - | | 0.41 | 2050 | 0.0076 | - | | 0.415 | 2075 | 0.0034 | - | | 0.42 | 2100 | 0.0072 | - | | 0.425 | 2125 | 0.0055 | - | | 0.43 | 2150 | 0.0086 | - | | 0.435 | 2175 | 0.0062 | - | | 0.44 | 2200 | 0.0036 | - | | 0.445 | 2225 | 0.0061 | - | | 0.45 | 2250 | 0.0108 | - | | 0.455 | 2275 | 0.0049 | - | | 0.46 | 2300 | 0.0079 | - | | 0.465 | 2325 | 0.0036 | - | | 0.47 | 2350 | 0.0042 | - | | 0.475 | 2375 | 0.0072 | - | | 0.48 | 2400 | 0.0103 | - | | 0.485 | 2425 | 0.0041 | - | | 0.49 | 2450 | 0.0048 | - | | 0.495 | 2475 | 0.0061 | - | | 0.5 | 2500 | 0.0043 | 0.9821 | | 0.505 | 2525 | 0.0118 | - | | 0.51 | 2550 | 0.0078 | - | | 0.515 | 2575 | 0.0071 | - | | 0.52 | 2600 | 0.0064 | - | | 0.525 | 2625 | 0.0048 | - | | 0.53 | 2650 | 0.0053 | - | | 0.535 | 2675 | 0.0058 | - | | 0.54 | 2700 | 0.0042 | - | | 0.545 | 2725 | 0.0057 | - | | 0.55 | 2750 | 0.0073 | - | | 0.555 | 2775 | 0.0040 | - | | 0.56 | 2800 | 0.0052 | - | | 0.565 | 2825 | 0.0052 | - | | 0.57 | 2850 | 0.0049 | - | | 0.575 | 2875 | 0.0037 | - | | 0.58 | 2900 | 0.0047 | - | | 0.585 | 2925 | 0.0042 | - | | 0.59 | 2950 | 0.0066 | - | | 0.595 | 2975 | 0.0058 | - | | 0.6 | 3000 | 0.0067 | 0.9864 |
### Training Time - **Training**: 5.1 hours - **Evaluation**: 13.4 minutes - **Total**: 5.4 hours ### Framework Versions - Python: 3.11.13 - Sentence Transformers: 5.7.0.dev0 - Transformers: 5.14.1 - PyTorch: 2.11.0+cu128 - Accelerate: 1.5.2 - Datasets: 3.5.0 - Tokenizers: 0.22.2 ## Additional Resources - [Sentence Transformers Documentation](https://www.sbert.net): the full documentation site, including training, evaluation, and pre-trained model catalogs. - [PyLate](https://github.com/lightonai/pylate): the upstream library whose features were absorbed into Sentence Transformers for multi-vector / late-interaction models. ## Citation ### BibTeX #### Sentence Transformers ```bibtex @inproceedings{reimers-2019-sentence-bert, title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2019", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/1908.10084", } ``` #### CachedMultiVectorMultipleNegativesRankingLoss ```bibtex @misc{gao2021scaling, title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup}, author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan}, year={2021}, eprint={2101.06983}, archivePrefix={arXiv}, primaryClass={cs.LG} } ```