Ribo-seq signal model: per-nucleotide ribosome profiling from sequence and RNA-seq
Predicts a per-nucleotide ribosome P-site profile for a transcript from its sequence and matched RNA-seq coverage. It takes any transcript, coding or non-coding, up to 10,000 nt, the longest in training. No ribosome-profiling experiment is required at inference.
The predicted profile can be fed to an ORF caller in place of real Ribo-seq, which is how it is evaluated here and what it is for: calling translated ORFs, including upstream and non-canonical ones, in samples where no Ribo-seq exists.
Two checkpoints are released. They differ only in the sequence mixer; body, heads, inputs, training data and ORF track are identical.
| checkpoint | mixer | params | device | profile r, 3-seed mean (range) |
|---|---|---|---|---|
mamba4_best.pt (primary) |
dilated CNN + 4 bidirectional Mamba blocks | 7,521,026 | GPU; CPU through the repository's pure-PyTorch reference | 0.6799 (0.6753-0.6851) |
attn_best.pt |
dilated CNN + 2 transformer layers | 5,071,106 | CPU or GPU | 0.6595 (0.6585-0.6603) |
Profile r is the median per-transcript Pearson correlation on raw counts over 70,883 held-out
transcripts, averaged across three training seeds. The released checkpoints are the seed-0 runs:
0.6851 for mamba4 and 0.6585 for attn, the values in <arch>_test_metrics.json.
Inputs
Ten channels per nucleotide:
| channels | content |
|---|---|
| 4 | one-hot A/C/G/T of the transcript |
| 5 | ORF-candidate track: frame 0/1/2 occupancy, start context, stop |
| 1 | RNA-seq coverage, depth-normalised |
Two output heads: a per-nucleotide profile (a distribution summing to 1) and a per-transcript count.
What must match training
The inputs are built from your own transcripts and RNA-seq with the repository's pack builder; the tutorial at https://ribosignal.readthedocs.io does this for chromosome 22. Two settings must match what the checkpoints were trained on:
- the ORF-candidate track is built with
--mode ext --kozak none. Both released models are--kozak nonemodels; feeding them a heuristic-Kozak track is a mistake that has already happened once in this project. - transcripts longer than 10,000 nt are left out. None was in training, and attention memory grows with the square of the length.
Reproducing the held-out metrics above additionally needs the training universe (84,472 transcripts), which is not published here.
Loading
import torch, json
cfg = json.load(open("mamba4_config.json"))
sd = torch.load("mamba4_best.pt", map_location="cpu") # plain state_dict
model.py is the model definition. With the copy here, mamba4_best.pt requires mamba_ssm
and a CUDA device. The repository's scripts/model.py falls back to a pure-PyTorch reference
(scripts/mamba_ref.py) where mamba_ssm is absent, so mamba4 also runs on CPU, more slowly.
attn_best.pt has no such requirement.
Training
- Human tissue data from Chothani et al.: Ribo-seq from GEO GSE182371, RNA-seq from GSE182372 (SuperSeries GSE182377).
- Leave-one-tissue-out: held out Hepatocytes; trained on Fibroblast, VSMC, ES, Fat, HA_EC, HCAEC, HUVEC. Brain dropped. 42,000 train / 8,400 val / 70,883 test transcripts.
- Training universe: 84,472 transcripts, protein-coding (74,704) and lncRNA (9,768), up to 10,000 nt. Other biotypes can be run but were not seen in training.
- Unique-mapper alignments for both assays (
--outFilterMultimapNmax 1); rRNA, tRNA, miRNA and mitochondrial loci removed before P-site calling. Mitochondrial protein-coding genes excluded throughout. - 256 channels, 10 blocks, dropout 0.1, 16,000 nt budget,
count_weight0.1,input_mode: both. - Checkpoint selected on validation
pearson_median.
License
MIT. Free for any use including commercial, with no attribution requirement beyond keeping the
copyright notice. Full text in LICENSE.
The weights and model.py are covered. The training data are not: they are third-party public
datasets (GEO GSE182371 and GSE182372) under their own terms, and this licence makes no claim
about them.
Code
https://github.com/ericmalekos/ribosignal
model.py here is an earlier copy of that repository's scripts/model.py; the repository version
adds the CPU fallback for mamba4. The pack builder, prediction, ORF-calling and scoring pipeline
live there.
Citation
Manuscript in preparation. Until then cite this repository and https://github.com/ericmalekos/ribosignal.