DrugSpace-eval-8B
DrugSpace-eval-8B is the evaluation-oriented standalone DrugSpace embedding
model for English drug descriptions. Its Stage I MNTP model was trained on the
pre-2020 PubTator split, after which it received Stage II contrastive alignment
on OpenFDA drug-label text.
The Stage II LoRA adapter has already been merged into the Stage I model. This repository therefore contains a complete model and does not require a separate base model or PEFT adapter at inference time.
Model lineage
| Stage | Model or data |
|---|---|
| Foundation model | meta-llama/Meta-Llama-3.1-8B-Instruct |
| Stage I | MNTP adaptation on pre-2020 PubTator text: z-cao/DrugSpace-mntp-eval-8B |
| Stage II | OpenFDA contrastive-alignment adapter: z-cao/DrugSpace-openfda-lora-eval |
| Final release | Stage II adapter merged into the Stage I evaluation model |
Usage
Install LLM2Vec and its dependencies:
pip install llm2vec transformers accelerate
Load the merged model directly:
import torch
import torch.nn.functional as F
from llm2vec import LLM2Vec
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
model = LLM2Vec.from_pretrained(
base_model_name_or_path="z-cao/DrugSpace-eval-8B",
enable_bidirectional=True,
pooling_mode="mean",
max_length=512,
doc_max_length=400,
skip_instruction=True,
device_map=device,
torch_dtype=dtype,
)
texts = [
"Metformin is an oral biguanide medication used to improve glycemic control in type 2 diabetes.",
"Another complete English description of a drug.",
]
embeddings = model.encode(
texts,
batch_size=2,
show_progress_bar=True,
convert_to_tensor=True,
)
similarity = F.cosine_similarity(
embeddings[0].unsqueeze(0),
embeddings[1].unsqueeze(0),
).item()
print(embeddings.shape)
print(similarity)
The released configuration uses mean pooling, a maximum sequence length of
512, a document maximum length of 400, and skip_instruction=true.
Model details
- Model type: bidirectional Llama-based LLM2Vec embedding model
- Parameters: approximately 8B
- Weight dtype: bfloat16
- Task: drug-description embedding, semantic similarity, and retrieval
- Language: English
- Output: one fixed-size embedding per input description
Training data
Stage I: pre-2020 PubTator
The Stage I evaluation model was adapted with an MNTP objective on the pre-2020 PubTator split.
Stage II: OpenFDA
Stage II used OpenFDA-derived single-ingredient labeling records. Text views were drawn from the following label sections when available:
- indications and usage
- clinical pharmacology
- mechanism of action
- pharmacodynamics
- pharmacokinetics
- contraindications
- warnings and cautions
- drug interactions
- adverse reactions
The processed source contained 1,867 records. Of these, 1,857 records with at least two distinct non-empty text views were eligible for contrastive training.
Stage II training procedure
- Fixed training set constructed once with seed 42 and reused for all 20 epochs
- 3,714 constructed triplets; 3,712 used per epoch after dropping the final incomplete batch
- Global batch size 64
- Learning rate 1e-4
- 300 warm-up steps
- Maximum sequence length 512
- bfloat16 training
- Final training step used as the released checkpoint; no alignment validation split or early stopping
Merge details
The Stage II adapter was merged into z-cao/DrugSpace-mntp-eval-8B with
PEFT merge_and_unload(safe_merge=True) and saved as bfloat16 safetensors. The
saved standalone checkpoint was reloaded successfully and checked against the
unmerged base-plus-adapter pipeline.
Intended use
This model is intended for evaluation-oriented research involving English drug
descriptions, including embedding generation, semantic similarity, clustering,
and retrieval. Use z-cao/DrugSpace-8B for the main non-evaluation release.
Limitations
- Only the Stage I PubTator training data follows the pre-2020 split; Stage II OpenFDA data was not temporally split.
- The model is intended for representation learning, not text generation.
- Training data and label-section availability may introduce coverage and documentation biases.
- Performance may degrade for non-English text, very short names, incomplete descriptions, or text outside the biomedical drug domain.
- Similarity in the embedding space does not establish therapeutic equivalence, safety, efficacy, or causal relationships.
- The model is for research use and must not be used as a substitute for professional medical judgment.
Main release
For the main DrugSpace release, use
z-cao/DrugSpace-8B.
License
Use of this model is subject to the applicable Llama 3.1 license and acceptable-use terms.
- Downloads last month
- 1
Model tree for z-cao/DrugSpace-eval-8B
Base model
meta-llama/Llama-3.1-8B