SentenceTransformer based on sentence-transformers/all-distilroberta-v1

This is a sentence-transformers model finetuned from sentence-transformers/all-distilroberta-v1 on the bio-embedding-finetuning dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'RobertaModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
  (2): Normalize({})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("johnnySoftData/distilroberta-ai-bio-embeddings")
# Run inference
queries = [
    'By providing an accessible, quantitative method to study complex cell communities, cytoNet helps researchers better understand environmental effects on cellular behavior.',
]
documents = [
    'We introduce cytoNet, a cloud-based tool to characterize cell populations from microscopy images. cytoNet quantifies spatial topology and functional relationships in cell communities using principles of network science. Capturing multicellular dynamics through graph features, cytoNet also evaluates the effect of cell-cell interactions on individual cell phenotypes. We demonstrate cytoNet’s capabilities in four case studies: 1) characterizing the temporal dynamics of neural progenitor cell communities during neural differentiation, 2) identifying communities of pain-sensing neurons in vivo, 3) capturing the effect of cell community on endothelial cell morphology, and 4) investigating the effect of laminin α4 on perivascular niches in adipose tissue. The analytical framework introduced here can be used to study the dynamics of complex cell communities in a quantitative manner, leading to a deeper understanding of environmental effects on cellular behavior. The versatile, cloud-based format of cytoNet makes the image analysis framework accessible to researchers across domains.',
    'Hand, foot and mouth disease caused by enterovirus 71(EV71) leads to the majority of neurological complications and death in young children. While putative inactivated vaccines are only now undergoing clinical trials, no specific treatment options exist yet. Ideally, EV71 specific intravenous immunoglobulins could be developed for targeted treatment of severe cases. To date, only a single universally neutralizing monoclonal antibody against a conserved linear epitope of VP1 has been identified. Other enteroviruses have been shown to possess major conformational neutralizing epitopes on both the VP2 and VP3 capsid proteins. Hence, we attempted to isolate such neutralizing antibodies against conformational epitopes for their potential in the treatment of infection as well as differential diagnosis and vaccine optimization. Here we describe a universal neutralizing monoclonal antibody that recognizes a conserved conformational epitope of EV71 which was mapped using escape mutants. Eight escape mutants from different subgenogroups (A, B2, B4, C2, C4) were rescued; they harbored three essential mutations either at amino acid positions 59, 62 or 67 of the VP3 protein which are all situated in the “knob” region. The escape mutant phenotype could be mimicked by incorporating these mutations into reverse genetically engineered viruses showing that P59L, A62D, A62P and E67D abolish both monoclonal antibody binding and neutralization activity. This is the first conformational neutralization epitope mapped on VP3 for EV71.',
    'Coordination of growth between and within organs contributes to the generation of well-proportioned organs and functionally integrated adults. The mechanisms that help to coordinate the growth between different organs start to be unravelled. However, whether an organ is able to respond in a coordinated manner to local variations in growth caused by developmental or environmental stress and the nature of the underlying molecular mechanisms that contribute to generating well-proportioned adult organs under these circumstances remain largely unknown. By reducing the growth rates of defined territories in the developing wing primordium of Drosophila, we present evidence that the tissue responds as a whole and the adjacent cell populations decrease their growth and proliferation rates. This non-autonomous response occurs independently of where growth is affected, and it is functional all throughout development and contributes to generate well-proportioned adult structures. Strikingly, we underscore a central role of Drosophila p53 (dp53) and the apoptotic machinery in these processes. While activation of dp53 in the growth-depleted territory mediates the non-autonomous regulation of growth and proliferation rates, effector caspases have a unique role, downstream of dp53, in reducing proliferation rates in adjacent cell populations. These new findings indicate the existence of a stress response mechanism involved in the coordination of tissue growth between adjacent cell populations and that tissue size and cell cycle proliferation can be uncoupled and are independently and non-autonomously regulated by dp53.',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[ 0.7204, -0.1550,  0.0768]])

Evaluation

Metrics

Triplet

  • Datasets: ai-job-validation, ai-bio-validation and ai-bio-test
  • Evaluated with TripletEvaluator
Metric ai-job-validation ai-bio-validation ai-bio-test
cosine_accuracy 1.0 1.0 1.0

Training Details

Training Dataset

bio-embedding-finetuning

  • Dataset: bio-embedding-finetuning at 2353672
  • Size: 3,568 training samples
  • Columns: query, abstract, and abstract_neg
  • Approximate statistics based on the first 100 samples:
    query abstract abstract_neg
    type string string string
    modality text text text
    details
    • min: 21 tokens
    • mean: 30.87 tokens
    • max: 49 tokens
    • min: 46 tokens
    • mean: 284.75 tokens
    • max: 512 tokens
    • min: 42 tokens
    • mean: 306.94 tokens
    • max: 512 tokens
  • Samples:
    query abstract abstract_neg
    Computational modeling demonstrates that eukaryotic histones can form stable homodimers, but their disordered tails generally favor the formation of heterotypic dimers. Histones compact and store DNA in both Eukarya and Archaea, forming heterodimers in Eukarya and homodimers in Archaea. Despite this, the folding mechanism of histones across species remains unclear. Our study addresses this gap by investigating 11 types of histone and histone-like proteins across humans, Drosophila, and Archaea through multiscale molecular dynamics (MD) simulations, complemented by NMR and circular dichroism experiments. We confirm and elaborate on the widely applied “folding upon binding” mechanism of histone dimeric proteins and report a new alternative conformation, namely, the inverted non-native dimer, which may be a thermodynamically metastable configuration. Protein sequence analysis indicated that the inverted conformation arises from the hidden ancestral head-tail sequence symmetry underlying all histone proteins, which is congruent with the previously proposed histone evolution hypotheses. Finally, to explore the potential formations of homodimers in Eukarya,... BackgroundLeprosy is the most frequent treatable neuromuscular disease. Yet, every year, thousands of patients develop permanent peripheral nerve damage as a result of leprosy. Since early detection and treatment of neuropathy in leprosy has strong preventive potential, we conducted a cohort study to determine which test detects this neuropathy earliest.
    Analyzing the invasive potential of Wolbachia is difficult due to significant variations in climate, release protocols, and human density across different deployment sites. Background The symbiotic bacterium Wolbachia is currently being trialled as a biocontrol agent in several countries to reduce dengue transmission. Wolbachia can invade and spread to infect all individuals within wild mosquito populations, but requires a high rate of maternal transmission, strong cytoplasmic incompatibility and low fitness costs in the host in order to do so. Additionally, extensive differences in climate, field-release protocols, urbanization level and human density amongst the sites where this bacterium has been deployed have limited comparison and analysis of Wolbachia’s invasive potential. Hair follicle stem cells (HFSCs) are multipotent cells that cycle through quiescence and activation to continuously fuel the production of hair follicles. Prior genome mapping studies had shown that tri-methylation of histone H3 at lysine 27 (H3K27me3), the chromatin mark mediated by Polycomb Repressive Complex 2 (PRC2), is dynamic between quiescent and activated HFSCs, suggesting that transcriptional changes associated with H3K27me3 might be critical for proper HFSC function. However, functional in vivo studies elucidating the role of PRC2 in adult HFSCs are lacking. In this study, by using in vivo loss-of-function studies we show that, surprisingly, PRC2 plays a non-instructive role in adult HFSCs and loss of PRC2 in HFSCs does not lead to loss of HFSC quiescence or changes in cell identity. Interestingly, RNA-seq and immunofluorescence analyses of PRC2-null quiescent HFSCs revealed upregulation of genes associated with activated state of HFSCs. Altogether, our findings show that tra...
    The efficacy of intermediate state inhibitors critically depends on the final disposition and irreversible deactivation of the inhibitor-bound target states. Both equilibrium and nonequilibrium factors influence the efficacy of pharmaceutical agents that target intermediate states of biochemical reactions. We explored the intermediate state inhibition of gp41, part of the HIV-1 envelope glycoprotein complex (Env) that promotes viral entry through membrane fusion. This process involves a series of gp41 conformational changes coordinated by Env interactions with cellular CD4 and a chemokine receptor. In a kinetic window between CD4 binding and membrane fusion, the N- and C-terminal regions of the gp41 ectodomain become transiently susceptible to inhibitors that disrupt Env structural transitions. In this study, we sought to identify kinetic parameters that influence the antiviral potency of two such gp41 inhibitors, C37 and 5-Helix. Employing a series of C37 and 5-Helix variants, we investigated the physical properties of gp41 inhibition, including the ability of inhibitor-bound gp41 to recover its fusion activity once inhibitor was removed f... Composed of hundreds of microbial species, the composition of the human gut microbiota can vary with chronic diseases underlying health disparities that disproportionally affect ethnic minorities. However, the influence of ethnicity on the gut microbiota remains largely unexplored and lacks reproducible generalizations across studies. By distilling associations between ethnicity and differences in two US-based 16S gut microbiota data sets including 1,673 individuals, we report 12 microbial genera and families that reproducibly vary by ethnicity. Interestingly, a majority of these microbial taxa, including the most heritable bacterial family, Christensenellaceae, overlap with genetically associated taxa and form co-occurring clusters linked by similar fermentative and methanogenic metabolic processes. These results demonstrate recurrent associations between specific taxa in the gut microbiota and ethnicity, providing hypotheses for examining specific members of the gut microbiota as med...
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Evaluation Dataset

bio-embedding-finetuning

  • Dataset: bio-embedding-finetuning at 2353672
  • Size: 446 evaluation samples
  • Columns: query, abstract, and abstract_neg
  • Approximate statistics based on the first 100 samples:
    query abstract abstract_neg
    type string string string
    modality text text text
    details
    • min: 19 tokens
    • mean: 31.18 tokens
    • max: 45 tokens
    • min: 74 tokens
    • mean: 303.82 tokens
    • max: 512 tokens
    • min: 43 tokens
    • mean: 302.8 tokens
    • max: 512 tokens
  • Samples:
    query abstract abstract_neg
    Mutant studies reveal that spirochetes lacking the bloodstream protein Vmp suffer reduced survival and inability to relapse, while lacking the tick protein Vtp prevents infection by tick bite. Borrelia hermsii, a causative agent of relapsing fever of humans in western North America, is maintained in enzootic cycles that include small mammals and the tick vector Ornithodoros hermsi. In mammals, the spirochetes repeatedly evade the host’s acquired immune response by undergoing antigenic variation of the variable major proteins (Vmps) produced on their outer surface. This mechanism prolongs spirochete circulation in blood, which increases the potential for acquisition by fast-feeding ticks and therefore perpetuation of the spirochete in nature. Antigenic variation also underlies the relapsing disease observed when humans are infected. However, most spirochetes switch off the bloodstream Vmp and produce a different outer surface protein, the variable tick protein (Vtp), during persistent infection in the tick salivary glands. Thus the production of Vmps in mammalian blood versus Vtp in ticks is a dominant feature of the spirochete’s alternating life cycle. We constructed two mut... With the growth of artificial intelligence (AI)-ready datasets such as the National Health and Nutrition Examination Survey (NHANES), new opportunities for data-driven research are being created, but also generating risks of data exploitation by paper mills. In this work, we focus on two areas of potential concern for AI-supported research efforts. First, we describe the production of large numbers of formulaic single-factor analyses, relating single predictors to specific health conditions, where multifactorial approaches would be more appropriate. Employing AI-supported single-factor approaches removes context from research, fails to capture interactions, avoids false discovery correction, and is an approach that can easily be adopted by paper mills. Second, we identify risks of selective data usage, such as analyzing limited date ranges or cohort subsets without clear justification, suggestive of data dredging, and post-hoc hypothesis formation. Using a systematic literature search ...
    The study examines the combined effects of food insecurity, stunting, worm infections, and physical fitness on children's selective attention and academic achievement. Background Socioeconomically deprived children are at increased risk of ill-health associated with sedentary behavior, malnutrition, and helminth infection. The resulting reduced physical fitness, growth retardation, and impaired cognitive abilities may impede children’s capacity to pay attention. The present study examines how socioeconomic status (SES), parasitic worm infections, stunting, food insecurity, and physical fitness are associated with selective attention and academic achievement in school-aged children. Molecular profiling studies have shown that 85% of canine urothelial carcinomas (UC) harbor an activating BRAF V595E mutation, which is orthologous to the V600E variant found in several human cancer subtypes. In dogs, this mutation provides both a powerful diagnostic marker and a potential therapeutic target; however, due to their relative infrequency, the remaining 15% of cases remain understudied at the molecular level. We performed whole exome sequencing analysis of 28 canine urine sediments exhibiting the characteristic DNA copy number signatures of canine UC, in which the BRAF V595E mutation was undetected (UDV595E specimens). Among these we identified 13 specimens (46%) harboring short in-frame deletions within either BRAF exon 12 (7/28 cases) or MAP2K1 exons 2 or 3 (6/28 cases). Orthologous variants occur in several human cancer subtypes and confer structural changes to the protein product that are predictive of response to different classes of small molecule MAPK pathway inhibi...
    Synthetic allergen-derived peptides can safely down-regulate allergic inflammation without the high risk of triggering mast cells found in native allergens. Background Synthetic peptides, representing CD4+ T cell epitopes, derived from the primary sequence of allergen molecules have been used to down-regulate allergic inflammation in sensitised individuals. Treatment of allergic diseases with peptides may offer substantial advantages over treatment with native allergen molecules because of the reduced potential for cross-linking IgE bound to the surface of mast cells and basophils. Nuclear size correlates with cell size, but the mechanism by which this scaling is achieved is not known. Here we screen fission yeast gene deletion mutants to identify essential factors involved in this process. Our screen has identified 25 essential factors that alter nuclear size, and our analysis has implicated RNA processing and LINC complexes in nuclear size control. This study has revealed lower and more extreme higher nuclear size phenotypes and has identified global cellular processes and specific structural nuclear components important for nuclear size control.
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 16
  • num_train_epochs: 1
  • learning_rate: 2e-05
  • warmup_steps: 0.1
  • per_device_eval_batch_size: 16
  • batch_sampler: no_duplicates

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 16
  • num_train_epochs: 1
  • max_steps: -1
  • learning_rate: 2e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 16
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss Validation Loss ai-job-validation_cosine_accuracy ai-bio-validation_cosine_accuracy ai-bio-test_cosine_accuracy
-1 -1 - - 1.0 1.0 -
0.4484 100 0.0088 0.0058 - 1.0 -
0.8969 200 0.0087 0.0060 - 1.0 -
1.0 223 - 0.0059 - 1.0 -
-1 -1 - - - 1.0 1.0

Training Time

  • Training: 12.4 minutes

Framework Versions

  • Python: 3.12.13
  • Sentence Transformers: 5.6.0
  • Transformers: 5.13.1
  • PyTorch: 2.11.0+cu128
  • Accelerate: 1.14.0
  • Datasets: 4.0.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
-
Safetensors
Model size
82.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for johnnySoftData/distilroberta-ai-bio-embeddings

Finetuned
(54)
this model

Dataset used to train johnnySoftData/distilroberta-ai-bio-embeddings

Papers for johnnySoftData/distilroberta-ai-bio-embeddings

Evaluation results