Oligonucleotide toxicity models
Nine endpoint-specific regression ensembles trained using complete expert descriptors, fixed XGBoost capacity, and an adaptive target-scale rule. This repository contains 90 native XGBoost UBJSON files: two branches × five seeds × nine endpoints. They are full-data refits for inference, not the held-out members used to estimate performance in the manuscript.
Code and reproduction instructions: https://github.com/DLeader-Inc/oligo-toxicity
Licence by endpoint
These licences apply to different subsets, not a choice of interchangeable licences for the whole repository. In particular, the two TLR models are not offered for commercial use. See LICENSE.md.
| Biological endpoint | Directory | Training rows | Weight licence |
|---|---|---|---|
| HeLa caspase-3/7 | models/cytotoxicity |
768 | CC BY 4.0 |
| LNA neuronal calcium score | models/neurotoxicity_lna |
1,825 | CC BY 4.0 |
| MOE mouse ICV FOB | models/neurotoxicity_moe |
2,437 | Apache-2.0 |
| TLR8 potentiation | models/tlr7_immunotox |
192 | CC BY-NC 4.0 |
| TLR7 inhibition | models/tlr8_immunotox |
192 | CC BY-NC 4.0 |
| Mouse ALT | models/mouse_hepatic |
2,701 | CC BY 4.0 |
| Rat ALT | models/rat_hepatic |
714 | CC BY 4.0 |
| Mouse FOB | models/mouse_neuro |
2,696 | CC BY 4.0 |
| Rat FOB | models/rat_neuro |
1,779 | CC BY 4.0 |
Counts describe endpoint records, not unique molecules across all datasets. The TLR directory labels are legacy identifiers inherited from the data packaging. The canonical biological endpoint names above are authoritative.
Inputs and prediction
Use reproduce/predict.py from the code repository. Pass a HELM string for
one oligonucleotide, interpreted from 5′ to 3′, and the canonical endpoint
identifier from catalog.json. Mouse/rat ALT also require actual
dosage_mg_per_kg, num_doses and dosing_period_days in assay_context.
No transcript sequence or drug target is inferred if it is not provided.
No remote Python code or pickle deserialization is needed for these weights.
from huggingface_hub import snapshot_download
# Pin revision to the desired immutable commit from this repository's history.
snapshot_download("DLeader/oligo-toxicity", revision="<commit>", local_dir="models")
python reproduce/predict.py --models models --input query.json --output prediction.json
Branch A uses position-wise molecular encoding and positional interactions. Branch B uses sequence-pattern, chemical-summary and joint descriptors. The exact ordered columns are stored in each endpoint manifest. The three reporting descriptor groups are not chemically independent: chemical information is also encoded in positional and regional columns.
Each seed averages its two branch outputs at equal weights. The target
scale is identity if training labels contain negatives, log1p if nonnegative
labels have Fisher–Pearson skewness greater than 2, and identity otherwise.
For log1p models, expm1 is applied to each fused seed prediction before the
five predictions are averaged. The chosen scale and response units are in
the manifest/catalog. Do not compare numerical values across different assays.
The prediction tool also computes TreeSHAP, an additive decomposition of the fitted tree prediction into its baseline and feature contributions. These contributions are on the fitted model scale and are not causal effects or raw-scale chemical substitution effects.
Validation and appropriate use
The manuscript's paired comparison uses 14,450 evaluation records across 45 source-native grouped outer partitions. The five OligoGym endpoints use fixed-seed nucleobase-cluster holdouts and the four Atlas endpoints use earliest-patent GroupKFold. Macro Spearman is 0.602890 for the integrated system versus 0.509139 for the local paper-best source replay, with higher Spearman on all nine endpoints. This is retrospective generalization to training-unseen assigned groups. Canonical base sequences can overlap, and there is no independent external or prospective validation cohort.
Serialization was checked for all 90 files on one reference row per endpoint: converted and original predictions were equal. Nine full ensemble predictions were also checked in a fresh CPU Python environment. These are software regression tests, not additional performance evaluation.
Research use only. Predictions are assay-specific associations, not clinical toxicity predictions or advice to administer an oligonucleotide. Applicability is limited by the represented lengths, chemistries and assay conditions. Ranking performance does not guarantee accurate absolute responses. A model that predicts a benefit from a chemical edit does not establish that benefit experimentally. Full-data refits must not be reused as held-out evidence.
Data provenance
Original datasets are not redistributed in this weight repository.
- Papargyri et al. (2020), DOI 10.1016/j.omtn.2019.12.011.
- Hagedorn et al. (2022), DOI 10.1089/nat.2021.0071.
- Alharbi et al. (2020), DOI 10.1093/nar/gkaa523.
- MOE mouse patent collection and OligoGym packaged files: fixed OligoGym source.
- Fixed ASO Atlas 2 dataset.
The code repository records source downloads, checksums, label definitions, frozen predictions and descriptor provenance. Model licences do not confer rights to practise patented inventions or clear every possible downstream use.