Cold-Start Chemical–Gene Ranking — LoRA adapter (1B)

Built with Llama. This is a LoRA (r=64) adapter for meta-llama/Llama-3.2-1B-Instruct, from the paper "Cold-Start Link Prediction Needs a Ranking Readout" (GaLM 2026 @ CIKM). The fine-tuned model is read out as a length-normalized sequence log-probability scorer to rank candidate genes for a chemical–gene interaction query, under a cold-start (unseen-chemical) protocol.

What this is

  • Adapter only (~172 MB). The base model meta-llama/Llama-3.2-1B-Instruct is not included — download it from the Hugging Face Hub (subject to the Llama Community License).
  • Fine-tuned on a CTD-derived chemical–gene QA corpus (augmented sample-47 footing).
  • Part of a size ladder released with the paper — 1B 0.78 / 3B 0.86 / 8B 0.92 (cold-start, sampled hard-negative MRR, K=99, sample-47). Under this capacity-limited LoRA regime, the readout improves monotonically with size, which — together with the full-fine-tuning result where a 1B already reaches the ceiling — locates the operative axis at capacity, not scale.

This adapter

  • Cold-start sampled hard-negative MRR = 0.783 (ep9; sample-47 corpus, K=99).
  • LoRA config: r=64, alpha=64, target modules q/k/v/o/gate/up/down_proj.

Intended use & limitations

  • Research use only. Cold-start chemical–gene ranking (scoring), not free generation and not clinical decision-making. Absolute values are only comparable within the sample-47, LoRA footing (never cross-compared with the full-fine-tuning / gl47 headline numbers).

Usage (sketch)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

base_id = "meta-llama/Llama-3.2-1B-Instruct"
tok  = AutoTokenizer.from_pretrained("BioRel/coldstart-lora-1b")
bnb  = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_quant_type="nf4")
base = AutoModelForCausalLM.from_pretrained(base_id, quantization_config=bnb, device_map={"": 0})
lm   = PeftModel.from_pretrained(base, "BioRel/coldstart-lora-1b").eval()

# score a candidate gene g for query q by length-normalized log-prob of g given q
# s(q, g) = (1/|g|) * sum_t log P(g_t | q, g_<t)  ; rank genes by s.
# Full scorer: https://github.com/BioRel/relational-qa-coldstart  (src/06_baselines/lora_scorer.py)

Training data & license

  • Data: derived from the Comparative Toxicogenomics Database (CTD), https://ctdbase.org. The dataset is subject to CTD terms; users must download CTD data themselves (terms: https://ctdbase.org/about/legal.jsp). CTD may access this dataset for quality control purposes. Non-commercial / research use.
  • Base model: meta-llama/Llama-3.2-1B-Instruct — governed by the Llama Community License (llama3.2). You must accept Meta's license to download the base.
  • Adapter weights: released for research use, subject to the base-model and CTD terms above.

Citation

@inproceedings{kim2026coldstart,
  title     = {Cold-Start Link Prediction Needs a Ranking Readout},
  author    = {Kim, Yunha and Kim, Young-Hak and Jun, Tae Joon},
  booktitle = {Proceedings of the Workshop on Graph-Augmented LLMs (GaLM), co-located with CIKM},
  year      = {2026}
}

Please also cite CTD: A. P. Davis et al., Comparative Toxicogenomics Database (CTD): update 2021, Nucleic Acids Research, 2021. https://ctdbase.org

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BioRel/coldstart-lora-1b

Adapter
(674)
this model