YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows

Repository layout

pCoMole/
  train.py / evaluate.py          # Edit Flow train / val-loss eval
  configs/                        # GFP (protein) and SELFIES configs
  model/ logic/ flow_matching/    # Edit Flow implementation
  smiles_tokenizer/               # SMILES SPE + SELFIES vocab
  data/selfies/28k_mimetics/      # shipped SELFIES training split
  gfp/                            # GFP generation + pCoMole
    pcomole.py                    # official GFP editor
    generate.py
    ckpt/last_2.ckpt
    classifier_ckpt/best.pt
    FPredX/                       # excitation / brightness / emission models
  peptidomimetics/                # peptidomimetic generation + pCoMole
    pcomole.py
    generate.py
    objectives.py                 # PeptiVerse + Admetica/DeepDTAGen mix
    ckpt/SELFIES_EditFlows.ckpt
    ckpt/SMILES_BindEvaluator.ckpt
    ckpt/admetica/                # LD50, solubility, Caco-2, half-life
    ckpt/deepdtagen/

1. Environment

A CUDA GPU is strongly recommended.

git clone <this-repo-url> pCoMole
cd pCoMole
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

HuggingFace model downloads happen on first use (facebook/esm2_t33_650M_UR50D, aaronfeller/PeptideCLM-23M-all, ChemBERTa).

PeptiVerse (required for peptidomimetic pCoMole)

Peptidomimetic property oracles use the latest PeptiVerse SMILES predictors. Clone it next to this repo, or set PEPTIVERSE_ROOT:

# recommended: sibling of pCoMole
cd ..
git clone https://huggingface.co/ChatterjeeLab/PeptiVerse
cd PeptiVerse
pip install -r requirements.txt
# follow PeptiVerse README to download model weights

pCoMole looks for PeptiVerse in this order:

  1. $PEPTIVERSE_ROOT
  2. ../PeptiVerse relative to this repository

PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged.

MAFFT (required for GFP pCoMole)

GFP excitation, brightness, and emission oracles run through FPredX and need MAFFT on your PATH.

# conda
conda install -c bioconda mafft

# or point to a local binary
export MAFFT_PATH=/path/to/mafft

Confirm with mafft --version before running GFP editing.

2. Train, evaluate, and sample Edit Flows

Run every command from the repository root.

SELFIES peptidomimetic Edit Flow

Training data is shipped at data/selfies/28k_mimetics.

# train
python train.py --config configs/config_selfies.yaml
# optional: python train.py --config configs/config_selfies.yaml --wandb

# evaluate a checkpoint on the validation split
python evaluate.py \
  --config configs/config_selfies.yaml \
  --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt

# unconditional / seeded generation
python peptidomimetics/generate.py \
  --config configs/config_selfies.yaml \
  --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \
  --input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \
  --num_steps 30 \
  --num_samples 2 \
  --output_csv outputs/peptidomimetic_unconditional.csv

Or: bash peptidomimetics/scripts/train.sh and bash peptidomimetics/scripts/generate.sh.

GFP protein Edit Flow

A trained GFP checkpoint is shipped at gfp/ckpt/last_2.ckpt. The original GFP training arrows are not included. To retrain, put a HuggingFace load_from_disk dataset under data/gfp/{train,validation} (or edit configs/config_gfp.yaml).

python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt

python gfp/generate.py \
  --config configs/config_gfp.yaml \
  --ckpt gfp/ckpt/last_2.ckpt \
  --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \
  --num_steps 10 \
  --num_samples 8 \
  --output_csv outputs/gfp_unconditional.csv

3. pCoMole multi-objective editing

GFP

Official editor: gfp/pcomole.py (length + excitation + brightness, with GFP-classifier and emission constraints).

bash gfp/scripts/pcomole.sh

Equivalent command:

python gfp/pcomole.py \
  --root_dir gfp/FPredX \
  --config configs/config_gfp.yaml \
  --ckpt gfp/ckpt/last_2.ckpt \
  --num_steps 10 \
  --num_candidates 50 \
  --num_rollouts 10 \
  --objective_weights 3 1 1 \
  --output_file outputs/gfp_length_excitation_brightness.csv \
  --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS'

If the output filename contains length_excitation_brightness, length_brightness, length_excitation, or length, the matching objective subset is used. Any other name uses all three objectives.

Peptidomimetics

Official editor: peptidomimetics/pcomole.py.

Objectives (7 scores when --specificity is on):

  1. non-toxicity
  2. solubility
  3. permeability
  4. half-life
  5. affinity
  6. motif
  7. specificity

PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint.

# after PeptiVerse is cloned and its weights are downloaded
bash peptidomimetics/scripts/pcomole.sh

--target and --motifs are required for affinity and motif oracles. --specificity adds the seventh objective.

4. Included checkpoints

Path Role
gfp/ckpt/last_2.ckpt GFP Edit Flow
gfp/classifier_ckpt/best.pt GFP hard constraint
gfp/FPredX/{ex,bright,em}_model FPredX property models
peptidomimetics/ckpt/SELFIES_EditFlows.ckpt peptidomimetic Edit Flow
peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt motif / specificity
peptidomimetics/ckpt/admetica/*.ckpt Admetica blending oracles
peptidomimetics/ckpt/deepdtagen/ DeepDTAGen affinity

Together these are about 10 GB. HuggingFace uploads should use Git LFS for the .ckpt / .pt / .pth files.

5. Environment variables

Variable Purpose
PEPTIVERSE_ROOT PeptiVerse repo root (manifest + training_classifiers/)
MAFFT_PATH MAFFT binary, if it is not on PATH
GFP_CLASSIFIER_CKPT optional override for the GFP classifier

6. Notes

  • Run scripts from the repo root, or use the wrappers under */scripts/.
  • GFP FPredX will fail immediately if mafft is missing; that is expected.
  • Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded.
  • Cas9 is not part of this release.
  • Training writes to outputs/ by default.

Citation

If you use this code, please cite the pCoMole paper and PeptiVerse.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support