You need to agree to share your contact information to access this model

LOAM is released for non-commercial research use only (see LICENSE and USE_POLICY.md). Access is granted automatically; your Hugging Face username and e-mail address are shared with Soilytix when you accept.

Log in or Sign Up to review the conditions and access this model content.

LOAM: Foundation models for microbial genomics

LOAM is Soilytix's family of genomic foundation models, bringing the diversity of environmental microbiomes to DNA sequence modelling. The models support sequence scoring, representation learning, and DNA continuation through a shared Hugging Face interface, processing sequences at single-nucleotide resolution with up to 8,192 nucleotides of context.

Trained exclusively on genomes assembled from Oxford Nanopore long-read sequencing, LOAM draws on a 67.5-billion-base-pair corpus of 15,640 species-representative microbial genomes from Microflora Danica. The 4 models, spanning 25M to 624M parameters, provide representations for gene-essentiality and enzyme-function prediction, alongside likelihood-based scores for zero-shot variant-effect analysis. The larger variants achieve competitive performance against genomic models with at least 10 times as many parameters on the evaluated prokaryotic benchmarks.

Preprint: coming soon on bioRxiv.

Model variants

The LOAM family offers 4 variants that share a training recipe and differ in transformer depth and width.

Model Parameters Depth (# blocks) Width Val. loss bits/nt
LOAM-25M 25.5M 5 640 1.6288
LOAM-100M 102.9M 8 1024 1.5134
LOAM-340M 340.0M 12 1536 1.4030
LOAM-624M 624.2M 16 1792 1.3463

This repository contains LOAM-100M. Start with LOAM-25M for pipeline development or limited memory. LOAM-624M achieves the lowest validation loss in the series.

Lower is better. Downstream task performance also depends on the task and embedding layer.

Benchmark results

LOAM is competitive with established genomic language models across 3 biological tasks. On each task, the leading LOAM variant scores above every evaluated model of equal or smaller size.

LOAM compared with other genomic language models on 3 biological tasks

Panel A shows how each task reads the frozen model. Panels B (Gene essentiality), C (Enzyme function) and D (RNA variant effects) plot score against measured parameter count for every model evaluated, with the LOAM series joined by a dashed line. Higher is better in every panel; the metric and the axis differ between them.

  • Gene essentiality (B): From the coding sequence of a bacterial gene and the promoter context upstream of it, predict whether the cell can survive losing that gene. Labels come from knockout screens, and the test genomes belong to genera held out of training, so the question is whether indispensability is recoverable from representations learned without ever seeing an essentiality label. LOAM-340M reaches 0.7722 macro mean AUROC, compared with Evo 1.5 (8k, 7B) at 0.7779 with 19.0 times as many parameters and Evo2-7B at 0.7880 with 19.1 times as many parameters.
  • Enzyme function (C): From the DNA of a gene, predict the Enzyme Commission class of the enzyme it encodes. This is the DNA arm of DGEB, so the prediction is made from nucleotides rather than from the translated protein, and genes catalysing related reactions have to land near one another in representation space. LOAM-624M reaches 0.3939 F1, the highest score among all 13 models evaluated.
  • RNA variant effects (D): Rank single-nucleotide variants of a gene by their measured effect on fitness. The prokaryotic deep mutational scanning assays in RNAGym read out phenotypes such as antibiotic resistance, enzyme activity, molecular interactions and folding stability, each against quantitative experimental measurements. LOAM-624M reaches 0.3176 macro Spearman correlation, compared with Evo 1.5 (8k, 7B) at 0.3182 with 10.3 times as many parameters.

Classification results use linear probes on frozen embeddings and report the best tested layer. RNA variant scoring is zero-shot; scoring methods and context limits differ between models. Each result represents one checkpoint and one seed.

Get started

Accept the conditions on the model page, then set up an environment and sign in:

uv venv --python 3.11
source .venv/bin/activate
uv pip install --torch-backend=auto "torch>=2.11" "transformers>=5.17" huggingface_hub
hf auth login

This uses uv (pip install uv if you don't have it). It provides Python 3.11 even if your system Python is older, and --torch-backend=auto installs the PyTorch build that matches your GPU driver. The release was tested end to end with PyTorch 2.11.0 and Transformers 5.17.0.

LOAM uses Peri-LN, which normalizes both the inputs and outputs of each attention and feed-forward sublayer. This changes the transformer block structure, so stock backbones such as LlamaForCausalLM are not drop-in replacements. trust_remote_code=True loads the LOAM implementation included in this repository, preserving the architecture used during training.

Load the model and tokenizer once for all three workflows. This setup uses a GPU when available, or CPU. Change repo_id to use another variant.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "Soilytix/LOAM-100M"
device = "cuda" if torch.cuda.is_available() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo_id, trust_remote_code=True, dtype=torch.float32,
).to(device).eval()

Before using your own sequences, read Tokenizer and input preparation.

Score DNA sequences

Compute mean next-nucleotide loss in bits. Lower values mean the sequence is more predictable to the model; they do not establish biological function.

import math

sequence = "ATGAGTAAAGGAGAAGAACTTTTCACTGGAGTTGTC"
inputs = tokenizer(sequence, return_tensors="pt", return_token_type_ids=False).to(device)

with torch.inference_mode():
    loss = model(**inputs, labels=inputs["input_ids"], use_cache=False).loss

print(loss.item() / math.log(2), "bits per nucleotide")

The first base is excluded because it has no preceding context. Use at least two bases. Keep float32 for close likelihood comparisons. For padded batches, follow the batching guidance in Tokenizer and input preparation.

Extract embeddings for biological prediction

Get one vector per nucleotide, then average over real bases to obtain one vector per sequence.

batch = tokenizer(
    ["ACGTTGCA", "GGTACCAATG"], padding=True,
    return_tensors="pt", return_token_type_ids=False,
).to(device)

with torch.inference_mode():
    outputs = model(**batch, output_hidden_states=True, use_cache=False)

per_nucleotide = outputs.hidden_states[-1]
mask = batch["attention_mask"].unsqueeze(-1)
per_sequence = (per_nucleotide * mask).sum(dim=1) / mask.sum(dim=1)
print(per_sequence.shape)  # [number of sequences, embedding width]

This uses the final normalized state. For downstream prediction, compare embedding layers on your validation set: the last layer is not always the most informative.

Generate DNA continuations

Continue a DNA sequence from a prompt. Sampling is recommended; greedy decoding can become repetitive.

tokenizer.padding_side = "left"  # Required when generating in padded batches.
prompt = "ATGAGTAAAGGAGAAGAACTTTTCACTGGAGTTGTC"
inputs = tokenizer(prompt, return_tensors="pt", return_token_type_ids=False).to(device)

with torch.inference_mode():
    generated = model.generate(
        **inputs, max_new_tokens=1000, do_sample=True, temperature=1.0,
    )

continuation = generated[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(continuation, skip_special_tokens=True))

Tokenizer and input preparation

All LOAM variants share a single-nucleotide tokenizer: one base, one token. Its vocabulary contains 58 token IDs, but only uppercase A, C, G, T are supported biological inputs for these weights. The remaining IDs cover ambiguity codes, soft masking, special tokens and reserved slots. Being accepted by the tokenizer does not mean a token was learned.

Input Tokenizer behavior How to prepare it
A, C, G, T One trained token per base Pass directly.
Lowercase letters Each becomes REPEAT, losing its base identity Use .upper() if you intend to discard soft masking.
N, R, Y, S, W, K, M, B, D, H, V Distinct tokens with untrained embeddings Split into unambiguous segments; do not join across unknown bases.
U Becomes UNK For RNA, use .upper().replace("U", "T") before tokenizing.
Whitespace, gaps, digits and unsupported symbols Become UNK silently Parse FASTA first, remove headers, join sequence lines within each record, and reject remaining non-ACGT characters.
  • Sequence boundaries: The tokenizer adds no start or end token. Do not insert BOS, EOS or other special tokens yourself. Generate from a nonempty DNA prompt and set max_new_tokens; EOS was not learned as a stopping signal or a biological boundary.
  • Length: Stay within 8,192 nucleotides, counting both prompt and continuation. The tokenizer does not truncate automatically. Split longer sequences into windows.
  • Batching: Use padding_side="left" for generation and pass attention_mask. Exclude padding when pooling embeddings. For batched loss, use padding_side="right" and set padded labels to -100, so they do not contribute to the score.
  • Strand: Both strands were represented through reverse-complement augmentation. A sequence and its reverse complement can still receive different scores.

Training and limitations

LOAM uses a decoder-only transformer trained to predict the next nucleotide. The source corpus contains 67.5 billion base pairs of dereplicated metagenome-assembled genomes from Microflora Danica. Each model was trained for approximately 1 epoch over the training split with reverse-complement augmentation. Whole genomes were assigned to separate splits; species-level separation was not established.

The training data reflects one environmental survey. Validate performance on your own taxa and tasks. Generated DNA requires experimental validation, and likelihood is not a measure of function or safety. LOAM takes DNA prompts and is not instruction-tuned.

Licence and responsible use

Code and tokenizer: Apache-2.0. The model weights are licensed under CC BY-NC-SA 4.0, the full text of which is in LICENSE. Corpus attribution is recorded in NOTICE.

Follow the Acceptable Use Policy. No pathogen screen, exclusion list or sequence-of-concern filter was applied to the training corpus. Prophage, plasmid and mobile-element sequence integrated into the MAGs may be present. Output intended for synthesis requires provider screening and institutional biosafety review.

Cite LOAM

If you use these models, please cite our preprint. It will be posted on bioRxiv shortly.

@article{ferry2026loam,
  title   = {{LOAM}: A family of genomic language models trained on long-read soil metagenomes},
  author  = {Ferry, Quentin RV. and Frank, Maurice and Steinkraus, Bruno R. and Rajakumar, Timothy},
  journal = {bioRxiv},
  year    = {2026},
  doi     = {10.1101/XXXX.XX.XX.XXXXXX},
  url     = {https://doi.org/10.1101/XXXX.XX.XX.XXXXXX},
  note    = {Preprint forthcoming on bioRxiv; DOI to be added once posted}
}

Acknowledgements

LOAM would not exist without Microflora Danica, the project that surveyed Danish environmental microbiomes and released the resulting genomes openly. We thank the teams behind the atlas and behind the long-read sequencing methodology it rests on (see NOTICE for full attribution).

  • Atlas. Singleton, C. M., Jensen, T. B. N., Delogu, F., et al. (2026). The Microflora Danica atlas of Danish environmental microbiomes. Nature 649:971-981. doi:10.1038/s41586-025-09794-2
  • Sequencing methodology. Sereika, M., Kirkegaard, R. H., Karst, S. M., et al. (2022). Oxford Nanopore R10.4 long-read sequencing enables the generation of near-finished bacterial genomes from pure cultures and metagenomes without short-read or reference polishing. Nature Methods 19(7):823-826. doi:10.1038/s41592-022-01539-7
Downloads last month
6
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Soilytix/LOAM-100M