Ribo-seq signal model: per-nucleotide ribosome profiling from sequence and RNA-seq

Predicts a per-nucleotide ribosome P-site profile for a transcript from its sequence and matched RNA-seq coverage. It takes any transcript, coding or non-coding, up to 10,000 nt, the longest in training. No ribosome-profiling experiment is required at inference.

The predicted profile can be fed to an ORF caller in place of real Ribo-seq, which is how it is evaluated here and what it is for: calling translated ORFs, including upstream and non-canonical ones, in samples where no Ribo-seq exists.

Two checkpoints are released. They differ only in the sequence mixer; body, heads, inputs, training data and ORF track are identical.

checkpoint mixer params device profile r, 3-seed mean (range)
mamba4_best.pt (primary) dilated CNN + 4 bidirectional Mamba blocks 7,521,026 GPU; CPU through the repository's pure-PyTorch reference 0.6799 (0.6753-0.6851)
attn_best.pt dilated CNN + 2 transformer layers 5,071,106 CPU or GPU 0.6595 (0.6585-0.6603)

Profile r is the median per-transcript Pearson correlation on raw counts over 70,883 held-out transcripts, averaged across three training seeds. The released checkpoints are the seed-0 runs: 0.6851 for mamba4 and 0.6585 for attn, the values in <arch>_test_metrics.json.

Inputs

Ten channels per nucleotide:

channels content
4 one-hot A/C/G/T of the transcript
5 ORF-candidate track: frame 0/1/2 occupancy, start context, stop
1 RNA-seq coverage, depth-normalised

Two output heads: a per-nucleotide profile (a distribution summing to 1) and a per-transcript count.

What must match training

The inputs are built from your own transcripts and RNA-seq with the repository's pack builder; the tutorial at https://ribosignal.readthedocs.io does this for chromosome 22. Two settings must match what the checkpoints were trained on:

  • the ORF-candidate track is built with --mode ext --kozak none. Both released models are --kozak none models; feeding them a heuristic-Kozak track is a mistake that has already happened once in this project.
  • transcripts longer than 10,000 nt are left out. None was in training, and attention memory grows with the square of the length.

Reproducing the held-out metrics above additionally needs the training universe (84,472 transcripts), which is not published here.

Loading

import torch, json

cfg = json.load(open("mamba4_config.json"))
sd  = torch.load("mamba4_best.pt", map_location="cpu")   # plain state_dict

model.py is the model definition. With the copy here, mamba4_best.pt requires mamba_ssm and a CUDA device. The repository's scripts/model.py falls back to a pure-PyTorch reference (scripts/mamba_ref.py) where mamba_ssm is absent, so mamba4 also runs on CPU, more slowly. attn_best.pt has no such requirement.

Training

  • Human tissue data from Chothani et al.: Ribo-seq from GEO GSE182371, RNA-seq from GSE182372 (SuperSeries GSE182377).
  • Leave-one-tissue-out: held out Hepatocytes; trained on Fibroblast, VSMC, ES, Fat, HA_EC, HCAEC, HUVEC. Brain dropped. 42,000 train / 8,400 val / 70,883 test transcripts.
  • Training universe: 84,472 transcripts, protein-coding (74,704) and lncRNA (9,768), up to 10,000 nt. Other biotypes can be run but were not seen in training.
  • Unique-mapper alignments for both assays (--outFilterMultimapNmax 1); rRNA, tRNA, miRNA and mitochondrial loci removed before P-site calling. Mitochondrial protein-coding genes excluded throughout.
  • 256 channels, 10 blocks, dropout 0.1, 16,000 nt budget, count_weight 0.1, input_mode: both.
  • Checkpoint selected on validation pearson_median.

License

MIT. Free for any use including commercial, with no attribution requirement beyond keeping the copyright notice. Full text in LICENSE.

The weights and model.py are covered. The training data are not: they are third-party public datasets (GEO GSE182371 and GSE182372) under their own terms, and this licence makes no claim about them.

Code

https://github.com/ericmalekos/ribosignal

model.py here is an earlier copy of that repository's scripts/model.py; the repository version adds the CPU fallback for mamba4. The pack builder, prediction, ORF-calling and scoring pipeline live there.

Citation

Manuscript in preparation. Until then cite this repository and https://github.com/ericmalekos/ribosignal.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support