Configuration Parsing Warning:In UNKNOWN_FILENAME: "auto_map.AutoTokenizer" must be a string

ESM-FineWeb-1B

OpenESM 1B model (variant d26) trained on FineWeb. This repository contains a Hugging Face Transformers-compatible export for the pretraining checkpoint. It uses OpenESM's custom architecture and remote model code; it is not the built-in transformers.EsmModel architecture.

How to Use

Install the runtime dependencies:

pip install torch transformers

Load the tokenizer and model with the custom OpenESM code enabled:

import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer

repo_id = "guan-wang/ESM-FineWeb-1B"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
)

device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
inputs = tokenizer("Hello world", return_tensors="pt").to(device)
with torch.inference_mode():
    outputs = model(**inputs)

print(outputs.logits.shape)

The first load may prompt you to review and allow the repository's custom model code. Loading remote Python code requires trust_remote_code=True.

Model Details

Detail Value
Model family OpenESM energy-based language model
Architecture Custom OpenESM implementation (d26)
Parameter scale 1B
Training dataset FineWeb
Training stage pretraining
Transformer blocks 26
Embedding dimension 1664
Attention heads 13
Context length 2048
Vocabulary size 32768

The parameter scale is the label used for this model in the OpenESM model configuration. The architecture and sequence settings above are read from the exported checkpoint configuration.

Files

  • config.json: model configuration.
  • model*.safetensors: model weights in the standard Transformers format.
  • modeling_esm.py: remote model and tokenizer implementation.
  • configuration_esm.py: remote configuration implementation.
  • tokenizer.pkl: serialized ESM tokenizer.
  • tokenizer_config.json: tokenizer auto-loading configuration.
  • token_bytes.pt: token byte table used by OpenESM metrics.

The original Lightning .ckpt is converted before upload and is not needed to load this repository. Training token counts are intentionally omitted from the repository name and file names.

Checkpoint Metadata

ebm-fineweb-d26-7b

EBM pretrained on FineWeb; d26, 7B.

This repository contains the following PyTorch checkpoint:

  • Archive filename: final-s=step=6999-d26-ctx2048.ckpt
  • Original filename: periodic-s=step=6999-d26-ctx2048.ckpt
  • Original relative path: pretrain/ebm/fineweb/d26/7B/periodic-s=step=6999-d26-ctx2048.ckpt

The checkpoint is uploaded as-is from the local training output. See the filename and training project for the exact architecture and loading code.

License

Please add the applicable model/data license before publishing this repo.

The code is maintained in the OpenESM repository.

Downloads last month
-
Safetensors
Model size
1B params
Tensor type
I64
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including guan-wang/ESM-FineWeb-1B