contacts-v1 tokenizer, <contacts-v1.multi> variant
The published contacts-v1 tokenizer with vocab id 7 renamed in place:
<contacts-and-distances-v1> -> <contacts-v1.multi>
Vocab size stays 2845 and every other id is unchanged, so a checkpoint
trained under the published tokenizer needs no embedding resize to be
fine-tuned under this one. Built by MarinFold exp230's make_multi_tokenizer.py
from the tokenizer shipped with contacts-v1-exp199-1.5B; the format is
MarinFold #163's
multi-draft document type.
It cannot read contacts-and-distances-v1 documents, and the published
tokenizer cannot read multi-draft documents. Ship it with the weights.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support