GPA checkpoints
Oracles and backbones trained for GPA: Generative Population Annealing for Test-Time Sequence Design with Pretrained Generative Models.
Code, run scripts and the training code that produced every file here: https://github.com/anirbansarkar-cs/gpa-paper-release
GPA itself trains nothing โ it is a test-time SMC sampler over a frozen generative prior. These are the oracles it designs against and the backbones it samples from, for the experiments that needed a model we trained rather than one from a published release.
Files
| File | What it is | Rebuilt by |
|---|---|---|
gpa_promoter_oracle_JURKAT.ckpt |
regLM EnformerModel activity oracle, JURKAT |
scripts/run_promoter_oracles.sh |
gpa_promoter_oracle_K562.ckpt |
same, K562 | scripts/run_promoter_oracles.sh |
gpa_promoter_oracle_THP1.ckpt |
same, THP1 | scripts/run_promoter_oracles.sh |
gpa_hyenadna_promoter_stage1.ckpt |
HyenaDNA fine-tuned as a conditional LM on promoter activity deciles | scripts/run_promoter_stage1.sh |
gpa_legnet_k562.ckpt |
LegNet K562 lentiMPRA activity oracle | model_zoo/lentimpra/mpralegnet.py |
gpa_legnet_k562_config.json |
config that mpralegnet.load_model needs beside the LegNet checkpoint |
โ |
The three promoter oracles and the stage-1 backbone are shared by both methods in the promoter comparison: GPA samples from the backbone frozen and designs against the oracles, and the Ctrl-DNA baseline RL-fine-tunes the same backbone against the same oracles.
Use
The fetch script in the code repository downloads these into the paths the loaders expect:
git clone https://github.com/anirbansarkar-cs/gpa-paper-release
cd gpa-paper-release
export GPA_DATA_ROOT=/path/to/datasets
bash scripts/fetch_checkpoints.sh
It renames them on the way in โ the promoter oracles become
human_paired_{jurkat,k562,THP1}.ckpt, which is the filename Ctrl-DNA's
base_optimizer.load_target_model looks for, and the backbone becomes
hyenadna_promoter_e3/best.ckpt.
To pull a single file directly:
from huggingface_hub import hf_hub_download
p = hf_hub_download("tataiani/gpa-checkpoints", "gpa_legnet_k562.ckpt")
Provenance and caveats
The promoter oracles and the stage-1 backbone were trained on a promoter MPRA built from Reddy et al., Designing Cell-Type-Specific Promoter Sequences Using Conservative Model-Based Optimization, NeurIPS 2024. The LegNet oracle was trained on lentiMPRA K562 data. Neither dataset is redistributed here or in the code repository; the code repository documents the input schema so you can rebuild them.
train_oracles.py sets no random seed, so retraining produces a different
checkpoint rather than a bit-identical one. These are the exact files used for
the reported runs.
The architecture for the promoter oracles is Ctrl-DNA's EnformerModel, imported
from their release rather than reimplemented โ loading them requires that code
on PYTHONPATH. See CTRL_DNA_HOME in the repository README.
Not included: our AlphaGenome-derived held-out evaluator. The code repository lists the public repositories you can use to fine-tune an equivalent one โ google-deepmind/alphagenome for the base model, MasayukiNagai/alphagenome-encoder-ft, Al-Murphy/alphagenome_FT_MPRA and genomicsxai/alphagenome_ft for the fine-tuning.
License
MIT, matching the code repository.