Manu-v0

Model description

Manu-v0 predicts the change in protein stability caused by a single amino-acid substitution.

The model accepts a wild-type protein sequence and a mutation such as V42A, then returns a predicted ฮ”ฮ”G value in kcal/mol.

Positive predicted ฮ”ฮ”G values indicate stabilization. Negative values indicate destabilization.

This release is the frozen Phase 6 deployment of the Cortex protein mutation stability predictor. It combines a frozen ESM-2 sequence encoder, an MLP regression head, and a small auxiliary residual correction selected on validation data.

Intended use

The model is intended for:

  • Research-oriented prioritization of single protein mutations.
  • Ranking candidate substitutions by predicted stability effect.
  • Exploring sequence-based protein engineering hypotheses.
  • Educational demonstrations of protein language model representations and regression.

The output is a computational estimate and should not be treated as an experimental measurement.

Input and output

Input

  • A canonical amino-acid sequence.
  • One single-substitution mutation in one-based notation, for example V42A.
  • Maximum supported sequence length: 1,024 residues.

The wild-type residue in the mutation must match the residue at the requested sequence position.

Output

  • Predicted ฮ”ฮ”G in kcal/mol.
  • Positive values mean predicted stabilization.
  • Negative values mean predicted destabilization.

Architecture

Protein sequence + single mutation
        โ†“
Input and mutation validation
        โ†“
Wild-type and mutant sequence construction
        โ†“
Frozen ESM-2 encoder
facebook/esm2_t12_35M_UR50D
        โ†“
Mutation-aware 3,840-dimensional representation
        โ†“
MLP regression head
3840 โ†’ 256 โ†’ 128 โ†’ 1
        โ†“
Auxiliary residual correction
        โ†“
Predicted ฮ”ฮ”G in kcal/mol

The ESM-2 encoder is frozen during regression training. The representation uses wild-type embeddings, mutant embeddings, their signed difference, absolute difference, and global sequence-level features.

The main MLP uses GELU activations, dropout of 0.1, Huber loss, and AdamW optimization. The Phase 6 residual blend weight is 0.13.

Evaluation

MegaScale cluster-held-out test

The primary evaluation uses a protein-cluster-held-out test split from MegaScale v2_230420. The test set contains 58,333 mutations from 63 proteins and 39 clusters.

Metric Result
MAE 0.63797 kcal/mol
RMSE 0.93231 kcal/mol
Pearson correlation 0.56772
Spearman correlation 0.57782
Bootstrap 95% interval for MAE 0.63267โ€“0.64352 kcal/mol

The Phase 6 blend was selected using validation data and evaluated once on the held-out test set.

External ThermoMutDB robustness check

As a secondary robustness check, the frozen deployment was evaluated on 5,345 strict sequence-independent unique variants from ThermoMutDB. This dataset contains heterogeneous experimental conditions and is not directly interchangeable with the MegaScale test score.

Metric Result
MAE 1.25277 kcal/mol
RMSE 1.86514 kcal/mol
Pearson correlation 0.37739
Spearman correlation 0.38869
Sign accuracy 72.09%

Training data

The model was trained using MegaScale release v2_230420.

  • Dataset: MegaScale
  • Zenodo DOI: 10.5281/zenodo.7992926
  • Split: megascale_cluster_split_v1
  • Split seed: 20260918
  • Target: experimental ฮ”ฮ”G in kcal/mol
  • Target convention: positive values represent stabilization

The Hugging Face dataset identifier listed in the metadata is LiteFold/MegaScale-Tsuboyama2023. Verify that it corresponds to the exact data release being redistributed with this model. The source dataset is licensed under CC-BY-4.0.

Base model

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for monabji/Manu-V0

Finetuned
(69)
this model

Dataset used to train monabji/Manu-V0

Evaluation results

  • Mean Absolute Error on MegaScale cluster-held-out test
    self-reported
    0.638
  • Root Mean Squared Error on MegaScale cluster-held-out test
    self-reported
    0.932
  • Pearson correlation on MegaScale cluster-held-out test
    self-reported
    0.568
  • Spearman correlation on MegaScale cluster-held-out test
    self-reported
    0.578