This model was asked to be published under my account, not the creators. The compute came from https://huggingface.co/posts/ProCreations/855858308074329

Boopit 1

A 27.3M-parameter language model with a 4096 context window and native BitNet b1.58 ternary weights (weights in {-1, 0, 1} from step 0).

This release contains the best held-out-loss checkpoint from pretraining on openbmb/Ultra-FineWeb-L1:

  • Selected checkpoint: step 3,400
  • Tokens seen by this checkpoint: 891,289,600
  • Validation loss: 4.480278

Training continued beyond this checkpoint, but later checkpoints did not improve held-out loss, so this best checkpoint was preserved and selected for release. The packed weights (model.boopit) are under 7 MB.

Architecture: 6 layers, width 512, 8 attention heads, tied embeddings, RMSNorm, GELU MLP, and a byte-level 16,384-token BPE vocabulary.

Chat fine-tune: ProCreations/boopit-1-chat.

Downloads last month
351
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ProCreations/boopit-1

Finetunes
1 model