GPA checkpoints

Oracles and backbones trained for GPA: Generative Population Annealing for Test-Time Sequence Design with Pretrained Generative Models.

Code, run scripts and the training code that produced every file here: https://github.com/anirbansarkar-cs/gpa-paper-release

GPA itself trains nothing โ€” it is a test-time SMC sampler over a frozen generative prior. These are the oracles it designs against and the backbones it samples from, for the experiments that needed a model we trained rather than one from a published release.

Files

File What it is Rebuilt by
gpa_promoter_oracle_JURKAT.ckpt regLM EnformerModel activity oracle, JURKAT scripts/run_promoter_oracles.sh
gpa_promoter_oracle_K562.ckpt same, K562 scripts/run_promoter_oracles.sh
gpa_promoter_oracle_THP1.ckpt same, THP1 scripts/run_promoter_oracles.sh
gpa_hyenadna_promoter_stage1.ckpt HyenaDNA fine-tuned as a conditional LM on promoter activity deciles scripts/run_promoter_stage1.sh
gpa_legnet_k562.ckpt LegNet K562 lentiMPRA activity oracle model_zoo/lentimpra/mpralegnet.py
gpa_legnet_k562_config.json config that mpralegnet.load_model needs beside the LegNet checkpoint โ€”

The three promoter oracles and the stage-1 backbone are shared by both methods in the promoter comparison: GPA samples from the backbone frozen and designs against the oracles, and the Ctrl-DNA baseline RL-fine-tunes the same backbone against the same oracles.

Use

The fetch script in the code repository downloads these into the paths the loaders expect:

git clone https://github.com/anirbansarkar-cs/gpa-paper-release
cd gpa-paper-release
export GPA_DATA_ROOT=/path/to/datasets
bash scripts/fetch_checkpoints.sh

It renames them on the way in โ€” the promoter oracles become human_paired_{jurkat,k562,THP1}.ckpt, which is the filename Ctrl-DNA's base_optimizer.load_target_model looks for, and the backbone becomes hyenadna_promoter_e3/best.ckpt.

To pull a single file directly:

from huggingface_hub import hf_hub_download
p = hf_hub_download("tataiani/gpa-checkpoints", "gpa_legnet_k562.ckpt")

Provenance and caveats

The promoter oracles and the stage-1 backbone were trained on a promoter MPRA built from Reddy et al., Designing Cell-Type-Specific Promoter Sequences Using Conservative Model-Based Optimization, NeurIPS 2024. The LegNet oracle was trained on lentiMPRA K562 data. Neither dataset is redistributed here or in the code repository; the code repository documents the input schema so you can rebuild them.

train_oracles.py sets no random seed, so retraining produces a different checkpoint rather than a bit-identical one. These are the exact files used for the reported runs.

The architecture for the promoter oracles is Ctrl-DNA's EnformerModel, imported from their release rather than reimplemented โ€” loading them requires that code on PYTHONPATH. See CTRL_DNA_HOME in the repository README.

Not included: our AlphaGenome-derived held-out evaluator. The code repository lists the public repositories you can use to fine-tune an equivalent one โ€” google-deepmind/alphagenome for the base model, MasayukiNagai/alphagenome-encoder-ft, Al-Murphy/alphagenome_FT_MPRA and genomicsxai/alphagenome_ft for the fine-tuning.

License

MIT, matching the code repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support