UNKAI
UNKAI is a protein-pair classification model for predicting whether two proteins are associated with the same enzymatic reaction.
The model takes per-residue protein embeddings as input, independently pools the two proteins using attention pooling, computes the absolute difference between the resulting protein-level representations, and performs binary classification with a multilayer perceptron.
Code and inference utilities are available at:
https://github.com/ukai3313/UNKAI
Model variants
This repository contains two pretrained UNKAI variants.
Original
Architecture:
Per-residue embeddings
|
Attention Pooling
|
Protein vectors v1, v2
|
|v1 - v2|
|
2560 -> 1599 -> 781 -> 117 -> 1
|
Sigmoid
The original model uses:
- Input embedding dimension: 2560
- Attention pooling
- Pair representation: absolute difference
|v1 - v2| - MLP dimensions:
2560 -> 1599 -> 781 -> 117 -> 1 - Dropout:
0.302766
Checkpoint:
original/model.pth
Seen-unseen
The seen-unseen model is designed for evaluation where one side of a validation or test pair belongs to a protein cluster observed during training, while the other side belongs to a cluster not observed during training.
Architecture:
Per-residue embeddings
|
Attention Pooling
with attention dropout
|
Protein vectors v1, v2
|
|v1 - v2|
|
2560 -> 1408 -> 640 -> 512 -> 1
The model uses:
- Input embedding dimension: 2560
- Attention pooling
- Attention dropout:
0.1627294925 - Pair representation: absolute difference
|v1 - v2| - MLP dimensions:
2560 -> 1408 -> 640 -> 512 -> 1 - Classifier dropout:
0.5282298379
Checkpoint:
seen_unseen/model.pth
The checkpoint produces logits. A sigmoid transformation is applied during inference to obtain a probability.
Input format
UNKAI expects one per-residue embedding file for each protein.
Expected filename:
<UNIPROT_ACCESSION>_embedding.npy
Supported array shapes are:
(L, 2560)
or:
(1, L, 2560)
where L is the protein sequence length.
The released checkpoints were trained using 2560-dimensional per-residue protein embeddings.
Installation
Clone the UNKAI GitHub repository:
git clone https://github.com/ukai3313/UNKAI.git
cd UNKAI
pip install -r requirements.txt
Add the source directory to PYTHONPATH:
export PYTHONPATH="$PWD/src"
Inference
Original model
python -m unkai.predict \
--model original \
--protein1 Q6GZV6 \
--protein2 Q6GZN7 \
--embeddings-dir /path/to/embeddings \
--checkpoint /path/to/original/model.pth
Seen-unseen model
python -m unkai.predict \
--model seen_unseen \
--protein1 Q6GZV6 \
--protein2 Q6GZN7 \
--embeddings-dir /path/to/embeddings \
--checkpoint /path/to/seen_unseen/model.pth
Example output:
Protein 1 : Q6GZV6
Protein 2 : Q6GZN7
Model : original
Probability : 0.002518
Prediction : 0
The current inference utility uses a probability threshold of 0.5.
Prediction = 1: predicted to be associated with the same enzymatic reaction.Prediction = 0: predicted not to be associated with the same enzymatic reaction.
Files
original/
โโโ model.pth
seen_unseen/
โโโ model.pth
Datasets
Datasets used with UNKAI are released separately:
ukaikotaro/UNKAI-dataset
Three dataset variants are provided:
- original
- seen_unseen
- strict
See the dataset repository for details about the splitting strategies.
Embeddings
The pretrained UNKAI checkpoints require precomputed per-residue protein embeddings.
The embedding collection used in this project is planned to be distributed separately.
Limitations
UNKAI predicts association between a pair of proteins with respect to enzymatic reactions and should not be interpreted as direct experimental evidence of identical biochemical function.
Performance may depend on the distribution of proteins, sequence similarity, clustering strategy, and embedding model used to construct the input representations.
The original and seen-unseen checkpoints were trained using different dataset splitting strategies and should therefore not be directly compared without considering the evaluation setting.
Citation
Citation information for the associated publication will be added here.
License
MIT License.