contacts-v1 tokenizer, <contacts-v1.multi> variant

The published contacts-v1 tokenizer with vocab id 7 renamed in place:

<contacts-and-distances-v1>  ->  <contacts-v1.multi>

Vocab size stays 2845 and every other id is unchanged, so a checkpoint trained under the published tokenizer needs no embedding resize to be fine-tuned under this one. Built by MarinFold exp230's make_multi_tokenizer.py from the tokenizer shipped with contacts-v1-exp199-1.5B; the format is MarinFold #163's multi-draft document type.

It cannot read contacts-and-distances-v1 documents, and the published tokenizer cannot read multi-draft documents. Ship it with the weights.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support