Configuration Parsing Warning:In UNKNOWN_FILENAME: "auto_map.AutoTokenizer" must be a string

CENO-1B-base

CENO-1B-base is the base pretraining checkpoint of the 1B CENO DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).

It is part of the CENO DNA foundation model family. Model code, the VEP pipeline, and a generation demo live in the companion CENO code repository. This checkpoint is standalone-loadable with trust_remote_code=True — the model code is bundled here.

Model details

Family CENO (base)
Training stage Base pretraining (stage 2)
Parameters 1.3B (1302.4M)
Precision bfloat16
model_type ceno
Architecture class CENOForCausalLM
Auto-map (model) modeling_ceno.CENOForCausalLM
Auto-map (tokenizer) ceno_tokenizer.CENOCharLevelTokenizer

Architecture

Property Value
Hidden layers 38
Hidden size 1024
Attention heads 16
Intermediate size 4096
Experts (MoE) 8 (top-2 per token)
Vocabulary 512 (byte / character-level)

The backbone is a Mamba / Attention / Mixture-of-Experts hybrid (Nemotron-H architecture). The tokenizer is character-level, mapping DNA bases to their ASCII byte codes.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

ckpt = "CladeTeam/CENO-1B-base"
model = AutoModelForCausalLM.from_pretrained(ckpt, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(ckpt, trust_remote_code=True)

ids = tokenizer.encode("ATCGATCG", return_tensors="pt")
# out = model.generate(ids, max_new_tokens=128)   # needs a CUDA GPU (Mamba kernels)

The Mamba layers require CUDA kernels, so forward passes and generation need a GPU. Config, tokenizer, and weight loading are CPU-safe.

Intended use

  • Base checkpoints (CENO-*) — genomic-sequence generation and embedding extraction; downstream adaptation (fine-tuning, probing) for genomics tasks.
  • MSA checkpoints (CENO-P-*) — variant effect prediction (VEP) by scoring wild-type vs. variant sequences with delta log-likelihood. See the TraitGym VEP example in the CENO code repository.

License

Apache-2.0. The bundled model code is derived from NVIDIA's Nemotron-H Hugging Face implementation (Apache-2.0); the tokenizer is derived from the Arc Institute Evo2 CharLevelTokenizer (Apache-2.0). See the LICENSE and NOTICE files in this repository for full attribution.

Downloads last month
413
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including CladeTeam/CENO-1B-base