DrugSpace-eval-8B

DrugSpace-eval-8B is the evaluation-oriented standalone DrugSpace embedding model for English drug descriptions. Its Stage I MNTP model was trained on the pre-2020 PubTator split, after which it received Stage II contrastive alignment on OpenFDA drug-label text.

The Stage II LoRA adapter has already been merged into the Stage I model. This repository therefore contains a complete model and does not require a separate base model or PEFT adapter at inference time.

Model lineage

Stage Model or data
Foundation model meta-llama/Meta-Llama-3.1-8B-Instruct
Stage I MNTP adaptation on pre-2020 PubTator text: z-cao/DrugSpace-mntp-eval-8B
Stage II OpenFDA contrastive-alignment adapter: z-cao/DrugSpace-openfda-lora-eval
Final release Stage II adapter merged into the Stage I evaluation model

Usage

Install LLM2Vec and its dependencies:

pip install llm2vec transformers accelerate

Load the merged model directly:

import torch
import torch.nn.functional as F
from llm2vec import LLM2Vec

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32

model = LLM2Vec.from_pretrained(
    base_model_name_or_path="z-cao/DrugSpace-eval-8B",
    enable_bidirectional=True,
    pooling_mode="mean",
    max_length=512,
    doc_max_length=400,
    skip_instruction=True,
    device_map=device,
    torch_dtype=dtype,
)

texts = [
    "Metformin is an oral biguanide medication used to improve glycemic control in type 2 diabetes.",
    "Another complete English description of a drug.",
]

embeddings = model.encode(
    texts,
    batch_size=2,
    show_progress_bar=True,
    convert_to_tensor=True,
)

similarity = F.cosine_similarity(
    embeddings[0].unsqueeze(0),
    embeddings[1].unsqueeze(0),
).item()

print(embeddings.shape)
print(similarity)

The released configuration uses mean pooling, a maximum sequence length of 512, a document maximum length of 400, and skip_instruction=true.

Model details

  • Model type: bidirectional Llama-based LLM2Vec embedding model
  • Parameters: approximately 8B
  • Weight dtype: bfloat16
  • Task: drug-description embedding, semantic similarity, and retrieval
  • Language: English
  • Output: one fixed-size embedding per input description

Training data

Stage I: pre-2020 PubTator

The Stage I evaluation model was adapted with an MNTP objective on the pre-2020 PubTator split.

Stage II: OpenFDA

Stage II used OpenFDA-derived single-ingredient labeling records. Text views were drawn from the following label sections when available:

  • indications and usage
  • clinical pharmacology
  • mechanism of action
  • pharmacodynamics
  • pharmacokinetics
  • contraindications
  • warnings and cautions
  • drug interactions
  • adverse reactions

The processed source contained 1,867 records. Of these, 1,857 records with at least two distinct non-empty text views were eligible for contrastive training.

Stage II training procedure

  • Fixed training set constructed once with seed 42 and reused for all 20 epochs
  • 3,714 constructed triplets; 3,712 used per epoch after dropping the final incomplete batch
  • Global batch size 64
  • Learning rate 1e-4
  • 300 warm-up steps
  • Maximum sequence length 512
  • bfloat16 training
  • Final training step used as the released checkpoint; no alignment validation split or early stopping

Merge details

The Stage II adapter was merged into z-cao/DrugSpace-mntp-eval-8B with PEFT merge_and_unload(safe_merge=True) and saved as bfloat16 safetensors. The saved standalone checkpoint was reloaded successfully and checked against the unmerged base-plus-adapter pipeline.

Intended use

This model is intended for evaluation-oriented research involving English drug descriptions, including embedding generation, semantic similarity, clustering, and retrieval. Use z-cao/DrugSpace-8B for the main non-evaluation release.

Limitations

  • Only the Stage I PubTator training data follows the pre-2020 split; Stage II OpenFDA data was not temporally split.
  • The model is intended for representation learning, not text generation.
  • Training data and label-section availability may introduce coverage and documentation biases.
  • Performance may degrade for non-English text, very short names, incomplete descriptions, or text outside the biomedical drug domain.
  • Similarity in the embedding space does not establish therapeutic equivalence, safety, efficacy, or causal relationships.
  • The model is for research use and must not be used as a substitute for professional medical judgment.

Main release

For the main DrugSpace release, use z-cao/DrugSpace-8B.

License

Use of this model is subject to the applicable Llama 3.1 license and acceptable-use terms.

Downloads last month
1
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for z-cao/DrugSpace-eval-8B

Finetuned
(1)
this model