RUMI-8M

RUMI-8M is a small decoder-only Transformer language model trained from scratch as part of the RUMI project.

Model Information

  • Model name: RUMI-8M
  • Parameters: 8,459,904
  • Architecture: Decoder-only Transformer
  • Vocabulary size: 16,000
  • Context length: 256
  • Embedding dimension: 192
  • Transformer layers: 12
  • Attention heads: 6
  • Feed-forward dimension: 768
  • Dropout: 0.1
  • Weight tying: enabled
  • Tokenizer: Byte-Level BPE

Important

RUMI-8M is a base language model.

It is not instruction-tuned and is not intended to behave like ChatGPT. Its primary capability is language-model text continuation.

Load RUMI-8M

Install the package dependencies:

pip install torch tokenizers huggingface_hub
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support