TopPDBLX construct boundary model
Give it a full-length protein sequence. It tells you, residue by residue, where a crystallographer would have cut.
Almost nobody crystallises the full-length gene product. You trim the flexible tails, drop a disordered linker, and try the folded core. Choosing where to cut is judgement, usually made once by hand from a disorder plot and an alignment.
The Protein Data Bank already holds 523,018 of those decisions, each made by someone who then got a structure. Every deposited chain records which stretch of the full-length protein was cloned, and SIFTS maps it back residue by residue. This model learns from them.
What it achieves
On 4,077 held-out proteins, split at 30% sequence identity so no homologue appears on both sides:
| Measure | Value | Why it matters |
|---|---|---|
| Boundary error | 9 residues (median) | How far the predicted cut sits from the real one, on the 1,314 test proteins that were genuinely truncated |
| MCC | 0.669 | A model that says "keep everything" scores 0.00 |
| Accuracy | 85.7% | Do not read alone. 61.5% of residues really are inside a construct, so "keep everything" already scores 61.5% |
Read the spread, not just the median
| boundary error | |
|---|---|
| Half the boundaries | within 5 residues |
| Three quarters | within 56 residues |
| Nine tenths | within 250 residues |
Excellent on most proteins, badly wrong on a minority. 60% of boundaries land within 10 residues and 56% of proteins have both ends within 25. The mean of 72 residues is the tail.
It half knows when it is wrong. Mean predicted probability across the span runs 0.97 on good predictions and 0.82 on bad ones. Gating at 0.85 covers 68% of proteins and lifts both-ends-within-25 from 56% to 71%. Treat a low-confidence span as a hint.
A worked example
Hen lysozyme (P00698) is 147 residues, of which 1 to 18 are the signal peptide and 19 to 147 the
mature chain. Asked cold, the model proposes 19 to 147. Nobody told it what a signal peptide
is; it learned that crystallographers do not clone them.
For scale: truncated constructs in this corpus trim a median of 62 residues from the N-terminus and keep under half the chain.
Usage
import torch, json
from transformers import AutoModel, AutoTokenizer
repo = "Dellboy/toppdblx-construct-boundary"
tok = AutoTokenizer.from_pretrained(repo)
esm = AutoModel.from_pretrained(repo)
head = torch.nn.Sequential(torch.nn.Dropout(0.1),
torch.nn.Linear(esm.config.hidden_size, 1))
head.load_state_dict(torch.load("boundary_head.pt")) # hf_hub_download this file
esm.eval(); head.eval()
sequence = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQ..." # your full-length protein
enc = tok(sequence, return_tensors="pt", truncation=True, max_length=1024)
with torch.no_grad():
prob = torch.sigmoid(head(esm(**enc).last_hidden_state).squeeze(-1))[0, 1:-1]
keep = (prob >= 0.5).nonzero().flatten()
print(f"suggested construct: residues {keep[0].item()+1} to {keep[-1].item()+1}")
Read this before trusting it
- It is not a disorder predictor. Residues that were cloned but never appeared in the density count as inside. It predicts what was cloned, not what turned out to be ordered.
- It only knows successes. Every label comes from a construct that produced a crystal. It has never seen one that failed, so it cannot say a boundary is bad, only that it is unlike the ones that worked.
- It errs towards keeping residues, predicting inside for 69% where truth is 61.5%. Proposed spans run slightly long. Trim rather than extend if choosing between it and your own judgement.
- Sequences longer than 1,022 residues are truncated by ESM-2's position limit.
Why this model and not a bigger one
This checkpoint has been challenged five ways and none of them beat it:
| Attempted | Outcome on the held-out test split |
|---|---|
| Six structural features: coil-avoidance, ESMFold pLDDT, Pfam domain edges, disorder prediction, disorder as a retrained input channel, surface entropy | None improved the boundary |
| Soft targets (train on the fraction of a protein's constructs covering each residue) | No gain, applied everywhere or gated to well-deposited proteins |
| Six epochs instead of three | No gain |
| ESM-2 t30-150M, four times the parameters | MCC 0.673 against 0.669: a dead heat, and worse on boundary error and every coverage@k |
The ceiling is in the labels, not the model. Where a crystallographer cuts carries real signal and this model extracts most of it, but the decision is also part convention, part whichever vector was to hand, and part arbitrary. None of that is recoverable from sequence, so a larger model simply fits the same ceiling more expensively.
Training
ESM-2 t12-35M fine-tuned end to end, 3 epochs, AdamW with one-cycle LR at 3e-4, batches bucketed by length into eight fixed widths. Labels are the per-residue consensus across every deposited construct for that protein, counted as inside where at least half the constructs include the residue. Validation MCC by epoch: 0.607, 0.681, 0.700, still improving at the end.
The failure condition (MCC below 0.40, or median boundary error worse than 20 residues) was declared before the run rather than after.
Dataset: TopPDBLX · Code: bellcheddar/TopPDBLX
- Downloads last month
- 12
Model tree for Dellboy/toppdblx-construct-boundary
Base model
facebook/esm2_t12_35M_UR50D