tomaarsen's picture
tomaarsen HF Staff
Add new MultiVectorEncoder model
e59c57b verified
|
Raw
History Blame Contribute Delete
62.1 kB
metadata
language:
  - en
license: apache-2.0
tags:
  - sentence-transformers
  - multi-vector
  - colbert
  - late-interaction
  - generated_from_trainer
  - dataset_size:640000
  - loss:CachedMultiVectorMultipleNegativesRankingLoss
widget:
  - text: >-
      This confirmed our hypothesis that the predictor variables would be
      different for men and women. Our results are similar to those of Knechtle
      et al. [23] who reported that, for male ironman athletes, anthropometric
      variables were important, as percent body fat was significantly associated
      with total race time. In female triathletes, training volume showed a
      relationship to total race time, in corroboration of our study.

       Another interesting finding was that the coefficient of determination of the models was higher in women (r 2 = 0.83) than in men (r 2 = 0.44). For women, the predicted race time did not correlate significantly (cor = 0.82, p-value = 0.09) with the achieved race time. For men, the predicted race time correlated significantly (cor = 0.84, p-value = 0.03) with the achieved race time. The differences in the coefficients of determination in the models might be explained by differences in anthropometric, training, and experience characteristics between women and men.

       A first important finding was that the personal best times in 5 km, 10 km, and half-marathon were the best predictors for female ultra-marathon performance. In the multiple regression models, the personal best half-marathon race time was significantly related to the ultra-marathon race time. Overall, it seems previous experience racing and fast personal best times are very important for ultra-marathon performance. This corroborates the results of Knechtle et al. [12] who examined 19 females in a 100 km ultra-marathon and found that the PBT in a marathon showed the highest correlation coefficient. Studies in other endurance sports disciplines such as triathlon showed personal best times in Olympic distance races were predictive in women for performance at Ironman distance [26] .

       Personal best times in marathon and in 5 km were associated with ultra-marathon race time for men. This adds to the bulk of knowledge available for males indicating that previous marathon personal best times seem to be a strong and independent predictor variable for ultra-endurance running performance in 100 km [15] , 350 km multi-day races [17] , and 24 h runs [17] . Previous studies have also shown that the personal best time in shorter races was also a predictor for Ironman race time in recreational male athletes [23, 27] , and PBT, not anthropometry or training volume, was associated with total race time in a triple-iron triathlon [24] . These findings of PBTs and high speed of running during training predicting ultramarathon performance reiterate the importance of intensity in training for men racing ultra-marathons.

       A recent study examined females racing in a 100 km distance ultra-marathon [12] . They found no association of race time with years running. We corroborate the results of Knechtle et al. [12] , as the variable years running was not associated with race time for women in this 62 km race. Rae et al. [28] examined the interaction of aging and racing on ultra-endurance running performance. Rae et al. [28] found that that overall athletes (18 women, 176 men) took approximately four years to reach peak running speed for a 56 km ultra-endurance race. It seemed that, regardless of the age at which the runners completed their first race, a period of about four years was required for the manifestation of adaptations associated with peak running performance during this ultra-endurance event. In our study, the average years running for females was 8.84, and so they were past this initial four years of improvement.

       Years running were associated to race time for males in bivariate analysis. This corroborates the results of Rae et al. [28] who studied mostly men (176 males, 18 females) and examined a similar distance to WUU2K (56 km versus 62 km of WUU2K). These findings contrast those of Knechtle et al. [17] , who examined multi-day racing male mountain runners, and of Knechtle et al. [15] , who examined 24 h race runners. For both of these studies, years running were not associated with ultra-marathon race time. Also, years running were not associated with marathon time for male marathoners [29] . Years running would seem to be more important for shorter runs, and maybe this would have to do with the intensity of training for performance in shorter runs, relying on less volume but more intense training, which would be easier for a non-novice runner.

       Recent studies show that age was an important predictor variable in ultra-marathon running [10] . Women's age was not significantly associated with race time in this 62 km race. This is in contrast to the results of Knechtle et al.
  - text: >-
      Being a resource and role model for their colleagues, nurse champions can
      contribute to improved quality of palliative care, when they have
      sufficient clinical experience, improved knowledge of palliative care,
      improved teaching capacities, and acquired authority towards managers and
      colleagues. [25, 28] . Still, rigorous evaluation of the effects of nurse
      champions on the outcome of care is necessary. In this article we describe
      the study protocol of the PalTeC-H project: a study on understanding and
      improving Palliative and Terminal Care in the Hospital by implementing a
      palliative care network of nurse champions.

       

       Objectives of this study are (1) to explore and understand the impact of the quality of care on the quality of life at the end of life and the quality of dying in a hospital and (2) to investigate the contribution of a quality improvement intervention which consists of the implementation of a network of palliative care nurse champions. We define end-of-life care as care provided during the last three days of life (at most). We hypothesize the implementation of the network to result in more attention for palliative care, in improved and timely recognition of patients' palliative care needs, in more involvement of palliative care experts and, eventually, in improved quality of life during the last three days of life, improved quality of dying and increased satisfaction of bereaved relatives.

       The intervention consists of the establishment of a palliative care network of nurse champions which indirectly affects care by three main components: education, knowledge dissemination and support, plus several organizational elements (Table 1) . On intervention wards two staff nurses are appointed to be palliative care nurse championsfurther referred to as champions. Together they form the palliative care network coordinated by the multidisciplinary consultation team for pain and palliative care. Champions participate in monthly educational meetings of the network and in a targeted education programme of two days annually. The education programme includes palliative care knowledge and skills as well as organizational knowledge and skills, e.g. on planning dissemination of knowledge, in order to teach the champions to be an ambassador of palliative care on the wards and a role model for their colleagues. The educational strategy is based on the principles of constructivist learning and includes multiple approaches [37] . A senior nurse consultant, member of the multidisciplinary consultation team, is assigned to be the network coordinator, supported by the medical oncologist of the team. This network coordinator facilitates the learning process of champions, by organizing meetings and education programmes, and supporting champions individually in their development and in performing activities. The monthly meetings stimulate the incremental grow of knowledge. Working and learning in a network throughout the hospital give champions the opportunity to share knowledge and learn from others' experiences, and to capture knowledge from outside their own working environment [23, 28, 34] .

       Champions need to identify gaps in knowledge on and quality of palliative care on their ward and to raise health care givers' awareness on patients' palliative care needs. They have to organize educational activities, implement protocols on palliative and terminal care, and evaluate these activities at the end of each year.

       Assuming that 14 champions each spend eight hours per month on network activities, and that the coordinator spends 24 hours per month, the intervention costs are estimated at  50.000 per year.

       All wards in a large general university hospital in the Netherlands participate in this study, including a specialized unit for palliative cancer care, but excluding the department of psychiatry and the Intensive Care departments.

       We collect data on adult patients who died at one of the 18 participating wards after having been admitted at least 6 hours prior to death.

       We designed a controlled before and after study with three phases: 1) pre-intervention phase (16 months); 2) phase in which the intervention is introduced (5 months); and 3) post-intervention phase (16 months). The intervention, i.e. the appointment of two champions joining the network, is introduced in seven wards that regularly admit cancer patients or patients with other chronic and life threatening diseases, such as chronic cardiac diseases and COPD. Although there is not much evidence on the time needed to effectively disseminate expertise and knowledge into clinical practice [31] [32] [33] , we decided that the introduction phase lasts five months, as a run-up period to generate gradual changes in champions' behavior [16, 38] . In the 11 wards where the intervention is not introduced, the same measurements are performed to control for changes that are not due to the intervention, for example changes in hospital policy (Table 2) .
  - text: >-
      I n medical care, treatment decisions made by clinicians and patients are
      generally based-implicitly or explicitly-on predictions of comparative
      outcome risks under alternative treatment conditions. Randomized
      controlled trials (RCTs), widely accepted as the gold standard for
      determining causal effects, have provided the primary evidence for these
      predictions. However, there is mounting recognition within evidence-based
      medicine of the limitations of RCTs as tools to guide clinical decision
      making at the individual patient level (1) (2) (3) (4) . Although
      historically the overall summary result from randomized trials ("average
      treatment effect") has been the cornerstone of evidence-based clinical
      decisions, interest is growing in understanding how a treatment's effect
      can vary across patients-a concept described as heterogeneity of treatment
      effects (HTE) (5) (6) (7) (8) (9) (10) (11) .

       Much literature exists on the limitations of conventional "1-variable-at-a-time" subgroup analyses, which serially divide the trial population into groups (for example, male vs. female or old vs. young) and examine the contrast in the treatment effect across these groups (12) (13) (14) (15) (16) (17) (18) (19) (20) (21) (22) . The limitations include risks for false-negative and false-positive results due to low power for statistical interactions, weak prior theory on potential effect modifiers, and multiplicity (4, 10, (23) (24) (25) . These analyses are also incongruent with the way clinical decision making occurs at the level of the individual patient, because patients have multiple attributes simultaneously that can affect the tradeoffs between the benefits and harms of the intervention. Individual patients thus belong to multiple subgroups, each of which may yield a different estimate of the treatment effect (4, 10) .

       The PATH (Predictive Approaches to Treatment effect Heterogeneity) Statement offers guidance relevant for "predictive" approaches to HTE analysis (26) that are designed to address some of the limitations mentioned in the previous paragraph. The goal of predictive HTE analysis is to provide individualized predictions of treatment effect, specifically defined by the difference between expected potential outcomes of interest with one intervention versus an alternative (4, 8) . We refer to this as the "individualized treatment effect." We avoid the term "individual treatment effects" because this latter term confusingly suggests that treatment effects can be estimated at the person level; such effects are inherently unobservable in parallel-group clinical trials because only 1 of 2 counterfactual potential outcomes can be observed (10, 27) . Individualized treatment effects have also been termed "conditional average treatment effects" (28) , denoting that they are the averaged treatment effect in a subpopulation (that is, conditioned on a set of covariates). However, for prediction, we are specifically interested in identifying the best conditional average treatment effect given all available patient characteristics, where "best" is defined as that which best discriminates between future patients who do and do not benefit from a treatment to optimize decision making for individual patients (29). By accounting for multiple variables simultaneously, predictive HTE analysis is foundational to the concept of personalization in evidence-based medicine (4) . Statement guidance focuses on identifying "clinically important HTE" (4, 7, 10) , or variation in the risk difference across patient subgroups that may be sufficient to span important decision thresholds that reflect treatment-related harms and burdens. The statement offers guidance on 2 distinct approaches to predictive HTE analysis (4) . With a "risk-modeling" approach, a multivariable model that predicts risk for an outcome (usually the primary study outcome) is first identified from external sources (an "external model") or developed directly on the trial population without a term for treatment assignment (an "internal model"). This prediction model is then applied to disaggregate patients within trials to examine risk-based variation in treatment effects. In a second approach, "effect modeling," a model is developed on RCT data with inclusion of a treatment assignment variable and potential inclusion of treatment interaction terms. These more flexible effect-modeling approaches have the potential to improve discrimination of patients who do and do not benefit, but they are especially vulnerable to overfitting and false discovery of promising subgroup effects (or they require very large databases that are well powered for the detection of interaction effects) (30). Both approaches can be used to predict individualized treatment effects-that is, the difference in expected outcome risks under 2 alternative treatments, conditional on important clinical variables. A fuller introduction to risk and effect modeling is presented in prior literature (4) .

       In this PATH Statement explanation and elaboration, we expand on the intent and motivation (and reservations) regarding the statements, criteria, considerations, and caveats.
  - text: >-
      Myocardial infarction (MI) remains the most frequent cardiovascular
      condition and can-beyond its immediate lethality-lead to cardiac failure
      and its associated late lethality. Cardiac failure is determined by the
      amount of myocardial tissue lost during ischaemia and, if reperfusion is
      achieved, the ensuing reperfusion injury, as well as by subsequent
      ventricular remodelling that adversely affects ventricular geometry.

       Ischaemic myocardial damage depends on cardiomyocyte apoptosis and necrosis. 1 Reperfusion injury is due to leucocyte-mediated cardiomyocyte bystander death during removal of necrosis. Growth of the defect also occurs secondary to stretch-induced tissue loss called non-ischaemic infarct expansion. 2, 3 Infarct healing can be divided into an early inflammatory and a late post-inflammatory phase. 4 The early inflammatory phase entails invasion of the infarcted tissue by leucocytes and removal of necrosis, population of the infarcted tissue by myofibroblasts and macrophages, and replacement of  These authors contributed equally to this work.

       necrosis by granulation tissue. The cellular events during this process in principle resemble wound healing in other tissues. 5 However, haemodynamic strain and the release of pro-hypertrophic growth factors in the myocardium during healing are responsible for an additional phenomenon: adverse remodelling that determines longterm functional outcome after infarction. 6 Syndecans are a family of transmembrane heparansulfate proteoglycans that regulate cell -cell and cell -matrix interactions. 7 Increased levels of syndecan-4 (Sdc4) were detected in the plasma of MI patients. 8 Expression of syndecan-1 (Sdc1) and Sdc4 is increased in the infarcted and the remote myocardium in animal models of MI. 9 Sdc1, which is mainly expressed in inflammatory and vascular cells, has recently been shown to affect ischaemic myocardial damage by reducing inflammation and thereby left ventricular (LV) dilatation after ischaemia. 10 Sdc4 is located within costamers and the Z-disc of cardiomyocytes, 11 which are thought to be important sites for mechano-sensing in cardiomyocytes. 12 Sdc4 has been shown to translate mechanical stretch into cytoplasmic signalling in fibroblasts. 13 Wound healing in the skin is disturbed and delayed in Sdc4-deficient mice, 14 but the role of Sdc4 for myocardial wound healing and early remodelling has not been elucidated. We therefore examined the effects of Sdc4 deficiency on myocardial damage and early infarct healing in mouse models of myocardial ischaemia and infarction.

       

       This study was approved by the Institutional Review Board and performed in accordance with the Guide for the Care and Use of Laboratory Animals published by the US National Institutes of Health. Sdc4 -/ -mice were backcrossed for more than 10 generations on C57BL/6 mice and ageand sex-matched Sdc4 -/ -and Sdc4 +/+ [wild-type (WT)] offspring of heterozygous matings were used for all studies. For myocardial ischaemia with reperfusion (MI/R) injury, the left coronary artery (LAD) was ligated for 30 min followed by 24 h of reperfusion. Area at risk (AAR) and infarct size were determined by TTC/Coomassie staining as described. 15 Data are presented as the average per cent infarct size per AAR. MI was induced by permanent ligation of the LAD as published previously. 16 Hearts were taken out 7 days later for molecular and histological analyses. Blood was collected from the retrobulbar plexus 24 h after MI or was drawn from the inferior vena cava before mice were sacrificed. Plasma was separated and troponin T levels were assessed by using Elecsys Troponin T high-sensitive test (Roche Diagnostics, Mannheim, Germany). 17 

       Primary neonatal rat ventricular cardiomyocytes were harvested from 1-to 3-day-old Sprague-Dawley rat pups (Charles River, Sulzfeld, Germany) as published elsewhere. 18 Cardiomyocytes were transfected with siRNA to Sdc4 (151941, Applied Biosystems, Darmstadt, Germany) or nonspecific (scr)-siRNA (1027281, Qiagen, Hilden, Germany) using RNAiMAX (Invitrogen, Karlsruhe, Germany).
  - text: >-
      ACR20 response rates were significantly higher with CZP plus MTX than
      placebo plus MTX at Week 1 (22.9 and 14.3% with CZP 200 mg plus MTX vs 5.6
      and 3.3% with placebo plus MTX in the RAPID 1 and 2 trials, respectively)
      [4, 5] . ACR20 response rates peaked at Week 12 in both studies (63.8 and
      62.7% for CZP 200 mg vs 18.3 and 12.7% for placebo in RAPID 1 and 2,
      respectively; both P < 0.001). At Week 24, ACR20 response rates were 58.8
      and 57.3% for patients receiving CZP 200 mg plus MTX, respectively, vs
      13.6 and 8.7%. The ITT populations for RAPID 1 and 2 consisted of all
      patients who were randomized into the studies; the modified ITT population
      for FAST4WARD consisted of all randomized patients who had taken one or
      more dose of study medication. Adapted from Mease [21] with permission of
      Future Medicine Ltd. CV: coefficient of variation; ITT:
      intention-to-treat; NA: not applicable.

       Significantly higher ACR50 and ACR70 response rates for CZP vs placebo groups were seen from Weeks 2 and 4 in RAPID 1, and Weeks 6 and 20 in RAPID 2, respectively. Responses were sustained to the end of the trials (Week 52 in RAPID 1 and Week 24 in RAPID 2; Table 2 ), and were similar in the CZP 400 mg plus MTX groups. CZP treatment also yielded significant improvements in all ACR core component scores, including reductions in swollen and tender joint scores and improvements in both patient's and physician's global assessments of disease activity, by Week 1 that were sustained throughout both studies [4, 5] . Treatment with CZP plus MTX was associated with significantly greater improvements in disease activity from Week 1, as evidenced by DAS-28 (ESR) scores, throughout both trials (P < 0.001 at all time points) [4, 5] . At Week 1, mean change from baseline in DAS-28 was À0.8 with CZP 200 mg and À0.3 with placebo in RAPID 1, and À0.8 with CZP 200 mg and À0.2 with placebo in RAPID 2. Improvements were sustained to the end of both studies (52 or 24 weeks, respectively; Fig. 1 ), and were similar with the CZP 400 mg dose. In RAPID 2, DAS-28 remission was observed in 9.4% of patients treated with CZP 200 mg plus MTX compared with only 0.8% of patients in the placebo group [5] .

       Both trials investigated the effects of CZP on the progression of joint damage. In RAPID 1, the mean (S.D.) change in mTSS from baseline to Week 52, which was a co-primary endpoint of the study, was significantly lower in patients receiving CZP 200 mg plus MTX [0.4 (5.7) in the CZP 200 mg group] compared with patients receiving placebo plus MTX [2.8 (7.8); P < 0.001] [4] . The changes were also significantly lower in the CZP plus MTX groups vs the placebo plus MTX group at Week 24 (P < 0.001). At both time points, significantly lower mean changes from baseline in both erosion (Week 24: 0 vs 0.7, Week 52: 0.1 vs 1.5; P < 0. [5] . Patients in the CZP 200 mg group in RAPID 2 also had significantly lower erosion (mean change from baseline: 0.1 vs 0.7) and joint space narrowing (mean change from baseline: 0.1 vs 0.5) subscores (P 4 0.01). Results for patients receiving the 400-mg dose were similar. An analysis of joint damage in patients who withdrew from the trials at Week 16 due to ACR20 non-response at Weeks 12 and 14 (as mandated by the study protocol) found that radiographic progression was inhibited by CZP plus MTX despite the fact that these patients did not meet the threshold for a clinical response [4, 5] .
pipeline_tag: feature-extraction
library_name: sentence-transformers
metrics:
  - maxsim_accuracy@1
  - maxsim_accuracy@3
  - maxsim_accuracy@5
  - maxsim_accuracy@10
  - maxsim_precision@1
  - maxsim_precision@3
  - maxsim_precision@5
  - maxsim_precision@10
  - maxsim_recall@1
  - maxsim_recall@3
  - maxsim_recall@5
  - maxsim_recall@10
  - maxsim_ndcg@10
  - maxsim_mrr@10
  - maxsim_map@100
model-index:
  - name: ColBERT gte-modernbert-base trained on MIRIAD question-passage pairs
    results:
      - task:
          type: multi-vector-information-retrieval
          name: Multi Vector Information Retrieval
        dataset:
          name: miriad eval
          type: miriad_eval
        metrics:
          - type: maxsim_accuracy@1
            value: 0.974
            name: Maxsim Accuracy@1
          - type: maxsim_accuracy@3
            value: 0.99
            name: Maxsim Accuracy@3
          - type: maxsim_accuracy@5
            value: 0.995
            name: Maxsim Accuracy@5
          - type: maxsim_accuracy@10
            value: 0.997
            name: Maxsim Accuracy@10
          - type: maxsim_precision@1
            value: 0.974
            name: Maxsim Precision@1
          - type: maxsim_precision@3
            value: 0.32999999999999996
            name: Maxsim Precision@3
          - type: maxsim_precision@5
            value: 0.19900000000000004
            name: Maxsim Precision@5
          - type: maxsim_precision@10
            value: 0.09970000000000001
            name: Maxsim Precision@10
          - type: maxsim_recall@1
            value: 0.974
            name: Maxsim Recall@1
          - type: maxsim_recall@3
            value: 0.99
            name: Maxsim Recall@3
          - type: maxsim_recall@5
            value: 0.995
            name: Maxsim Recall@5
          - type: maxsim_recall@10
            value: 0.997
            name: Maxsim Recall@10
          - type: maxsim_ndcg@10
            value: 0.9863523881462793
            name: Maxsim Ndcg@10
          - type: maxsim_mrr@10
            value: 0.9828250000000002
            name: Maxsim Mrr@10
          - type: maxsim_map@100
            value: 0.9830133116883116
            name: Maxsim Map@100

ColBERT gte-modernbert-base trained on MIRIAD question-passage pairs

This is a Multi-Vector Encoder model trained on the miriad-4.4_m-split dataset using the sentence-transformers library. It maps inputs to sequences of 128-dimensional token-level vectors and scores them with late interaction (MaxSim), useful for semantic search with late interaction.

Model Details

Model Description

  • Model Type: Multi-Vector Encoder
  • Maximum Sequence Length: 8192 tokens
  • Output Dimensionality: 128 dimensions
  • Similarity Function: maxsim
  • Supported Modality: Text
  • Training Dataset:
    • miriad-4.4_m-split
  • Language: en
  • License: apache-2.0

Model Sources

Full Model Architecture

MultiVectorEncoder(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'query_expansion': {'strategy': 'min', 'attend': False, 'token': None, 'length': 32}, 'architecture': 'ModernBertModel'})
  (1): Dense({'in_features': 768, 'out_features': 128, 'bias': False, 'activation_function': 'torch.nn.modules.linear.Identity', 'module_input_name': 'token_embeddings', 'module_output_name': 'token_embeddings'})
  (2): MultiVectorMask({'skiplist_words': ['!', '"', '#', '$', '%', '&', "'", '(', ')', '*', '+', ',', '-', '.', '/', ':', ';', '<', '=', '>', '?', '@', '[', '\\', ']', '^', '_', '`', '{', '|', '}', '~'], 'keep_only_token_ids': None})
  (3): Normalize({'module_input_name': 'token_embeddings', 'module_output_name': 'token_embeddings'})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import MultiVectorEncoder

# Download from the 🤗 Hub
model = MultiVectorEncoder("tomaarsen/multivector-gte-modernbert-base-miriad")
# Run inference: each input becomes a sequence of per-token vectors (variable length).
queries = [
    'How does treatment with CZP plus MTX impact disease activity in patients with rheumatoid arthritis?\n',
]
documents = [
    "ACR20 response rates were significantly higher with CZP plus MTX than placebo plus MTX at Week 1 (22.9 and 14.3% with CZP 200 mg plus MTX vs 5.6 and 3.3% with placebo plus MTX in the RAPID 1 and 2 trials, respectively) [4, 5] . ACR20 response rates peaked at Week 12 in both studies (63.8 and 62.7% for CZP 200 mg vs 18.3 and 12.7% for placebo in RAPID 1 and 2, respectively; both P < 0.001). At Week 24, ACR20 response rates were 58.8 and 57.3% for patients receiving CZP 200 mg plus MTX, respectively, vs 13.6 and 8.7%. The ITT populations for RAPID 1 and 2 consisted of all patients who were randomized into the studies; the modified ITT population for FAST4WARD consisted of all randomized patients who had taken one or more dose of study medication. Adapted from Mease [21] with permission of Future Medicine Ltd. CV: coefficient of variation; ITT: intention-to-treat; NA: not applicable.\n\n Significantly higher ACR50 and ACR70 response rates for CZP vs placebo groups were seen from Weeks 2 and 4 in RAPID 1, and Weeks 6 and 20 in RAPID 2, respectively. Responses were sustained to the end of the trials (Week 52 in RAPID 1 and Week 24 in RAPID 2; Table 2 ), and were similar in the CZP 400 mg plus MTX groups. CZP treatment also yielded significant improvements in all ACR core component scores, including reductions in swollen and tender joint scores and improvements in both patient's and physician's global assessments of disease activity, by Week 1 that were sustained throughout both studies [4, 5] . Treatment with CZP plus MTX was associated with significantly greater improvements in disease activity from Week 1, as evidenced by DAS-28 (ESR) scores, throughout both trials (P < 0.001 at all time points) [4, 5] . At Week 1, mean change from baseline in DAS-28 was À0.8 with CZP 200 mg and À0.3 with placebo in RAPID 1, and À0.8 with CZP 200 mg and À0.2 with placebo in RAPID 2. Improvements were sustained to the end of both studies (52 or 24 weeks, respectively; Fig. 1 ), and were similar with the CZP 400 mg dose. In RAPID 2, DAS-28 remission was observed in 9.4% of patients treated with CZP 200 mg plus MTX compared with only 0.8% of patients in the placebo group [5] .\n\n Both trials investigated the effects of CZP on the progression of joint damage. In RAPID 1, the mean (S.D.) change in mTSS from baseline to Week 52, which was a co-primary endpoint of the study, was significantly lower in patients receiving CZP 200 mg plus MTX [0.4 (5.7) in the CZP 200 mg group] compared with patients receiving placebo plus MTX [2.8 (7.8); P < 0.001] [4] . The changes were also significantly lower in the CZP plus MTX groups vs the placebo plus MTX group at Week 24 (P < 0.001). At both time points, significantly lower mean changes from baseline in both erosion (Week 24: 0 vs 0.7, Week 52: 0.1 vs 1.5; P < 0. [5] . Patients in the CZP 200 mg group in RAPID 2 also had significantly lower erosion (mean change from baseline: 0.1 vs 0.7) and joint space narrowing (mean change from baseline: 0.1 vs 0.5) subscores (P 4 0.01). Results for patients receiving the 400-mg dose were similar. An analysis of joint damage in patients who withdrew from the trials at Week 16 due to ACR20 non-response at Weeks 12 and 14 (as mandated by the study protocol) found that radiographic progression was inhibited by CZP plus MTX despite the fact that these patients did not meet the threshold for a clinical response [4, 5] .",
    'Five minutes after atropine, the R:T ratio increased from 1.15 (0.4) to 1.40 (0.6) (P < 0.01); at 30 min it was 1.51 (0.7) (P < 0.001) and at 60 min it was 1.33 (0.5) (P < 0.05). The R-wave amplitude was not affected by atropine. No changes in heart rate, QTc interval, RSA and R:T ratio occurred after placebo. COMMENT These data show that, in the presence of vagal block by atropine, the QTc interval increased significantly and the T-wave of the ECG was flattened.\n\n We chose a relatively large dose of atropine to ensure parasympathetic block as confirmed by the disappearance of RSA. Day, McComp and Campbell fl] have suggested that QT dispersion (interlead variability) gives an indication of arrythmogenicity and repolarization. We used a single lead V 2 which, according to the same group, provides the closest approximation to maximum QT interval [4] . They also accept the validity of a single lead value for QTc when changes are monitored. The flattened T-wave after atropine in our volunteers probably also reflected irregularity in repolarization.\n\n Atropine has been shown to increase the incidence of cardiac arrhythmia during induction of anaesthesia [3] . In addition, i.v. atropine has been shown to cause ventricular tachycardia in a patient with a prolonged QT interval syndrome [5] . Inhibition of the sympathoadrenal tone by opioids shortens the QTc interval in patients with vagal block. Vagal stimulation protects the heart against arrhythmogenic vulnerability [2] and against prolongation of the QT interval. In our study, the QTc interval was prolonged, probably because sympathoadrenal tone became dominant after parasympathetic block by atropine.\n\n In diabetic patients, vagal denervation develops gradually. Maintenance of remaining borderline vagal function by avoiding anticholinergics may be of value in diabetic patients, as serious cardiac arrhythmia has been described in these patients during anaesthesia and after atropine. Furthermore, ventricular fibrillation after i.v. atropine for bradycardia has been shown to occur in acute myocardial infarction [6] . The routine use of anticholinergics at induction of anaesthesia must be seriously questioned.',
    'This occurs since, in the folded state, the dansyl group is encapsulated in the hydrophobic cavity of the β-cyclodextrin ring resulting in a net fluorescence enhancement [99] . As a further development of this work, Riccardi and co-workers have described a tris-conjugated TBA 15 (tris-mTBA), equipped with a dansyl, a β-cyclodextrin and a biotin tag at the ends. This novel design has allowed the incorporation of TBA 15 onto streptavidin-coated NPs, leading to a remarkable increase of its anticoagulant properties. The developed systems have provided the basis for suitable aptamer-based devices for theranostic applications, allowing simultaneously both fluorescence-based detection and modulation of the thrombin activity [101] .\n\n Notably, in addition to the sensing approaches based on conformational switch random coil-G-quadruplex structure, also thrombin-induced changes starting from a hairpin structure are possible if the aptamer is properly engineered. In this context, Hamaguchi et al. have described a TBA 15 elongated at the 5 -end with few nucleotides complementary to the 3 -end and therefore able to adopt a stem-loop structure or hairpin [102] . In addition, the aptamer is equipped with a fluorescent/quencher pair, i.e., a fluorescein and a dabcyl moiety at the 5 -and 3 -end, respectively. In the absence of thrombin, the close proximity between the two reporter groups in the hairpin structure determines fluorescence quenching. After thrombin recognition, the stem-loop structure is destabilized in favour of interactions with the protein. Under these conditions, the fluorescent dye and the quencher are distant, thus allowing a "turn-on" of the fluorescence signal, indicative of the binding with the target molecule ( Figure 7c ).\n\n Alternative approaches for "structure switch signalling aptamers" are reported by Nutiu and Li [103] . Their strategy for designing aptamer-based fluorescent reporters involves structural switches from DNA/DNA duplex to DNA/target complex. In this study, the aptamer beacon consists of a tripartite duplex structure including a 5 -fluorescein-labeled oligomer (FDNA), a 3 -dabcyl-labeled oligomer (QDNA) and a longer oligonucleotide sequence comprising Stem-1 and Stem-2, complementary to FDNA and QDNA, respectively. Stem-2 also contains the TBA 15 sequence in a partial overhang (Figure 8a ). In the absence of the target protein, the aptamer naturally binds to FDNA and QDNA, bringing the fluorophore and the quencher in close proximity and thus completely inhibiting the fluorescence signal. The presence of thrombin triggers the formation of the aptamer-target complex, causing the release of QDNA and fully restoring the fluorescence emission.\n\n Cancers 2017, 9, 174 10 of 43\n\n Notably, in addition to the sensing approaches based on conformational switch random coil-G-quadruplex structure, also thrombin-induced changes starting from a hairpin structure are possible if the aptamer is properly engineered. In this context, Hamaguchi et al. have described a TBA15 elongated at the 5′-end with few nucleotides complementary to the 3′-end and therefore able to adopt a stem-loop structure or hairpin [102] . In addition, the aptamer is equipped with a fluorescent/quencher pair, i.e., a fluorescein and a dabcyl moiety at the 5′-and 3′-end, respectively. In the absence of thrombin, the close proximity between the two reporter groups in the hairpin structure determines fluorescence quenching. After thrombin recognition, the stem-loop structure is destabilized in favour of interactions with the protein. Under these conditions, the fluorescent dye and the quencher are distant, thus allowing a "turn-on" of the fluorescence signal, indicative of the binding with the target molecule ( Figure 7c ).\n\n Alternative approaches for "structure switch signalling aptamers" are reported by Nutiu and Li [103] . Their strategy for designing aptamer-based fluorescent reporters involves structural switches from DNA/DNA duplex to DNA/target complex. In this study, the aptamer beacon consists of a tripartite duplex structure including a 5′-fluorescein-labeled oligomer (FDNA), a 3′-dabcyl-labeled oligomer (QDNA) and a longer oligonucleotide sequence comprising Stem-1 and Stem-2, complementary to FDNA and QDNA, respectively. Stem-2 also contains the TBA15 sequence in a partial overhang (Figure 8a ).',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings[0].shape, document_embeddings[0].shape)
# (32, 128) (814, 128)

# Get the MaxSim similarity scores
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[26.4624,  9.1034,  3.9515]])

Evaluation

Metrics

Multi Vector Information Retrieval

Metric Value
maxsim_accuracy@1 0.974
maxsim_accuracy@3 0.99
maxsim_accuracy@5 0.995
maxsim_accuracy@10 0.997
maxsim_precision@1 0.974
maxsim_precision@3 0.33
maxsim_precision@5 0.199
maxsim_precision@10 0.0997
maxsim_recall@1 0.974
maxsim_recall@3 0.99
maxsim_recall@5 0.995
maxsim_recall@10 0.997
maxsim_ndcg@10 0.9864
maxsim_mrr@10 0.9828
maxsim_map@100 0.983

Training Details

Training Dataset

miriad-4.4_m-split

  • Dataset: miriad-4.4_m-split
  • Size: 640,000 training samples
  • Columns: question and passage_text
  • Approximate statistics based on the first 100 samples:
    question passage_text
    type string string
    modality text text
    details
    • min: 10 tokens
    • mean: 22.31 tokens
    • max: 60 tokens
    • min: 541 tokens
    • mean: 940.66 tokens
    • max: 1415 tokens
  • Samples:
    question passage_text
    What factors may contribute to increased pulmonary conduit durability in patients who undergo the Ross operation compared to those with right ventricular outflow tract obstruction?
    I n 1966, Ross and Somerville 1 reported the first use of an aortic homograft to establish right ventricle-to-pulmonary artery continuity in a patient with tetralogy of Fallot and pulmonary atresia. Since that time, pulmonary position homografts have been used in a variety of right-sided congenital heart lesions. Actuarial 5-year homograft survivals for cryopreserved homografts are reported to range between 55% and 94%, with the shortest durability noted in patients less than 2 years of age. 4 Pulmonary position homografts also are used to replace pulmonary autografts explanted to repair left-sided outflow disease (the Ross operation). Several factors may be likely to favor increased pulmonary conduit durability in Ross patients compared with those with right ventricular outflow tract obstruction, including later age at operation (allowing for larger homografts), more normal pulmonary artery architecture, absence of severe right ventricular hypertrophy, and more natural positioning of ...
    How does MCAM expression in hMSC affect the growth and maintenance of hematopoietic progenitors? After culture in a 3-dimensional hydrogel-based matrix, which constitutes hypoxic conditions, MCAM expression is lost. Concordantly, Tormin et al. demonstrated that MCAM is down-regulated under hypoxic conditions. 10 Furthermore, it was shown by others and our group that oxygen tension causes selective modification of hematopoietic cell and mesenchymal stromal cell interactions in co-culture systems as well as influence HSPC metabolism. [44] [45] [46] Thus, the observed differences between Sharma et al. and our data in HSPC supporting capacity of hMSC are likely due to the different culture conditions used. Further studies are required to clarify the influence of hypoxia in our model system. Altogether these findings provide further evidence for the importance of MCAM in supporting HSPC. Furthermore, previous reports have shown that MCAM is down-regulated in MSC after several passages as well as during aging and differentiation. 19, 47 Interestingly, MCAM overexpression in hMSC enhance...
    What is the relationship between Fanconi anemia and breast and ovarian cancer susceptibility genes?
    ( 31 ) , of which 5% -10 % may be caused by genetic factors ( 32 ) , up to half a million of these patients may be at risk of secondary hereditary neoplasms. The historic observation of twofold to fi vefold increased risks of cancers of the ovary, thyroid, and connective tissue after breast cancer ( 33 ) presaged the later syndromic association of these tumors with inherited mutations of BRCA1, BRCA2, PTEN, and p53 ( 16 ) . By far the largest cumulative risk of a secondary cancer in BRCA mutation carriers is associated with cancer in the contralateral breast, which may reach a risk of 29.5% at 10 years ( 34 ) . The Breast Cancer Linkage Consortium ( 35 , 36 ) also documented threefold to fi vefold increased risks of subsequent cancers of prostate, pancreas, gallbladder, stomach, skin (melanoma), and uterus in BRCA2 mutation carriers and twofold increased risks of prostate and pancreas cancer in BRCA1 mutation carriers; these results are based largely on self-reported family history inf...
  • Loss: CachedMultiVectorMultipleNegativesRankingLoss with these parameters:
    {
        "score_metric": "colbert_scores",
        "mini_batch_size": 8,
        "mini_batch_num_tokens": null,
        "score_mini_batch_size": 8,
        "scale": 1.0,
        "size_average": true,
        "gather_across_devices": false
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 128
  • num_train_epochs: 1
  • learning_rate: 3e-05
  • warmup_steps: 0.05
  • bf16: True
  • load_best_model_at_end: True
  • batch_sampler: no_duplicates
  • max_length: 1024

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 128
  • num_train_epochs: 1
  • max_steps: -1
  • learning_rate: 3e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.05
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: True
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 8
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}
  • max_length: 1024

Training Logs

Click to expand
Epoch Step Training Loss miriad_eval_maxsim_ndcg@10
-1 -1 - 0.9176
0.005 25 0.2384 -
0.01 50 0.1516 -
0.015 75 0.0738 -
0.02 100 0.0545 -
0.025 125 0.0356 -
0.03 150 0.0273 -
0.035 175 0.0282 -
0.04 200 0.0200 -
0.045 225 0.0218 -
0.05 250 0.0117 -
0.055 275 0.0121 -
0.06 300 0.0120 -
0.065 325 0.0137 -
0.07 350 0.0125 -
0.075 375 0.0154 -
0.08 400 0.0123 -
0.085 425 0.0100 -
0.09 450 0.0112 -
0.095 475 0.0109 -
0.1 500 0.0092 0.9788
0.105 525 0.0102 -
0.11 550 0.0105 -
0.115 575 0.0062 -
0.12 600 0.0113 -
0.125 625 0.0063 -
0.13 650 0.0118 -
0.135 675 0.0068 -
0.14 700 0.0086 -
0.145 725 0.0077 -
0.15 750 0.0103 -
0.155 775 0.0169 -
0.16 800 0.0097 -
0.165 825 0.0110 -
0.17 850 0.0072 -
0.175 875 0.0072 -
0.18 900 0.0081 -
0.185 925 0.0071 -
0.19 950 0.0091 -
0.195 975 0.0111 -
0.2 1000 0.0072 0.9794
0.205 1025 0.0068 -
0.21 1050 0.0076 -
0.215 1075 0.0076 -
0.22 1100 0.0076 -
0.225 1125 0.0146 -
0.23 1150 0.0068 -
0.235 1175 0.0062 -
0.24 1200 0.0094 -
0.245 1225 0.0063 -
0.25 1250 0.0103 -
0.255 1275 0.0070 -
0.26 1300 0.0075 -
0.265 1325 0.0072 -
0.27 1350 0.0053 -
0.275 1375 0.0043 -
0.28 1400 0.0091 -
0.285 1425 0.0092 -
0.29 1450 0.0077 -
0.295 1475 0.0092 -
0.3 1500 0.0064 0.9766
0.305 1525 0.0069 -
0.31 1550 0.0069 -
0.315 1575 0.0061 -
0.32 1600 0.0070 -
0.325 1625 0.0074 -
0.33 1650 0.0059 -
0.335 1675 0.0069 -
0.34 1700 0.0071 -
0.345 1725 0.0056 -
0.35 1750 0.0082 -
0.355 1775 0.0059 -
0.36 1800 0.0059 -
0.365 1825 0.0072 -
0.37 1850 0.0073 -
0.375 1875 0.0037 -
0.38 1900 0.0072 -
0.385 1925 0.0045 -
0.39 1950 0.0055 -
0.395 1975 0.0062 -
0.4 2000 0.0059 0.9781
0.405 2025 0.0057 -
0.41 2050 0.0076 -
0.415 2075 0.0034 -
0.42 2100 0.0072 -
0.425 2125 0.0055 -
0.43 2150 0.0086 -
0.435 2175 0.0062 -
0.44 2200 0.0036 -
0.445 2225 0.0061 -
0.45 2250 0.0108 -
0.455 2275 0.0049 -
0.46 2300 0.0079 -
0.465 2325 0.0036 -
0.47 2350 0.0042 -
0.475 2375 0.0072 -
0.48 2400 0.0103 -
0.485 2425 0.0041 -
0.49 2450 0.0048 -
0.495 2475 0.0061 -
0.5 2500 0.0043 0.9821
0.505 2525 0.0118 -
0.51 2550 0.0078 -
0.515 2575 0.0071 -
0.52 2600 0.0064 -
0.525 2625 0.0048 -
0.53 2650 0.0053 -
0.535 2675 0.0058 -
0.54 2700 0.0042 -
0.545 2725 0.0057 -
0.55 2750 0.0073 -
0.555 2775 0.0040 -
0.56 2800 0.0052 -
0.565 2825 0.0052 -
0.57 2850 0.0049 -
0.575 2875 0.0037 -
0.58 2900 0.0047 -
0.585 2925 0.0042 -
0.59 2950 0.0066 -
0.595 2975 0.0058 -
0.6 3000 0.0067 0.9864

Training Time

  • Training: 5.1 hours
  • Evaluation: 13.4 minutes
  • Total: 5.4 hours

Framework Versions

  • Python: 3.11.13
  • Sentence Transformers: 5.7.0.dev0
  • Transformers: 5.14.1
  • PyTorch: 2.11.0+cu128
  • Accelerate: 1.5.2
  • Datasets: 3.5.0
  • Tokenizers: 0.22.2

Additional Resources

  • Sentence Transformers Documentation: the full documentation site, including training, evaluation, and pre-trained model catalogs.
  • PyLate: the upstream library whose features were absorbed into Sentence Transformers for multi-vector / late-interaction models.

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

CachedMultiVectorMultipleNegativesRankingLoss

@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}